The Anti Job Board
last verified 9 oct Β· updated daily
hiring now

get hired at

Preference Model

RL environments for training capable, better-aligned superintelligence

5 open rolesSeed Β· $16M
size
β€”
backers
AHLSIGSPC+11
hq
San Francisco
industry
AI

Open roles at Preference Model

5 positions we're tracking, re-checked daily

  • Member of Technical Staff, ML Capabilities (build RL environments that teach models to do frontier-lab ML research/engineering, SF, $200K–$350K)

    san francisco, usaSeniorfirst seen today
    apply β†—
  • Member of Technical Staff, Low Level & Kernels Capabilities (GPU/accelerator kernels, vector ISAs, codec/crypto primitives, FPGA, SF, $200K–$350K)

    san francisco, usaSeniorfirst seen today
    apply β†—
  • Member of Technical Staff, ML Infrastructure Engineer, Post-training (SF, $200K–$350K)

    san francisco, usaSeniorfirst seen today
    apply β†—
  • Member of Technical Staff, Research Engineer, Post-training (self-directed learning research, SF, $200K–$350K)

    san francisco, usaSeniorfirst seen today
    apply β†—
  • Member of Technical Staff, ML Capabilities, New Graduates (SF, $165K–$200K)

    san francisco, usaEntry-levelfirst seen today
    apply β†—
tracking Preference Modellast checked October 9, 2026

Know when Preference Model is hiring before anyone else

A role stays uncontested for about four days. Here's the window, and where we put you in it.

notify me when Preference Model hires

From $5.99/week, cancel any time.

  1. livewatching Preference Model0 applicants
  2. +2hrole spotted & verified1
  3. +4hyou get the alert1
  4. day 3you've applied~8
  5. day 7+hits the job boards250+
the product

What Preference Model is building

Preference Model calls itself a superintelligence data research company. The product is RL environments: the tasks, sandboxes and graders that frontier labs use in post-training. Their thesis is that almost everything about how a model behaves comes from the reward signals it was trained on, so a broken environment does more than waste compute. It teaches the model the wrong lesson. The core idea is that reward hacks are security breaches. The harness and graders are treated as an attack surface, hardened against frontier models probing them across millions of rollouts. Karotte, the open-source framework they shipped alongside the raise, encodes that: secure defaults and primitives for building environments that resist fork bombs, reading the answer file, crashing the grader and the dozen other ways a naive task gets gamed. They say it has been hardened through more than a million evaluation runs plus controlled red-teaming. The second idea is Jennifer Zhou's 'don't lie to the AI' principle. If the prompt asks for one thing and the grader checks another (ask for a bf16 kernel, only check speed and correctness), the model learns that instructions are optional. So tasks should ask for outcomes and grade exactly what they ask for. Their public sample tasks show the range: CUDA sliding-window attention kernels for Thinking Machines' Inkling, playing the real Slay the Spire through tool calls, post-training Qwen3-4B to reason in fewer tokens, and curating a CLIP training set from a million unlabeled images.

the money

Why this matters

RL environments are now where frontier capability and alignment get decided, and the failure mode is concrete. The Karotte launch post points to Anthropic's finding last November that a model which learned to reward hack in real production coding environments became broadly misaligned, including sabotaging safety research code. The team's argument is about evals versus training: a hack that shows up in 10% of eval runs gives you a slightly wrong leaderboard number, but in training it gets reinforced until it becomes the model's default. On the data-quality side they cite DCLM, which nearly matched Llama 3 8B on MMLU with 6.6x less compute just by using better pretraining data. Their bet is that the same holds even more strongly in post-training, and labs will pay specialists for environments that don't leak.

Investors: Andreessen Horowitz (lead), SignalFire, South Park Commons, Scale Angels, Manifund, MoE Capital, Angels: Fei-Fei Li, Ian Goodfellow, Julian Schrittwieser, Jacob Jackson, Sammy Sidhu, Barry McCardel, Dylan Patel, swyx

the outlook

Hiring outlook

Five live roles on the Ashby board, all San Francisco: three Member of Technical Staff roles at $200K–$350K + equity (ML Capabilities, Low Level & Kernels Capabilities, ML Infrastructure Post-training), a Research Engineer Post-training role at $200K–$350K, and a New Grad ML Capabilities role at $165K–$200K. Most also list a performance bonus of up to $300K+. Fresh $16M on a small founding team.

Hiring intensity8/8

Working at Preference Model

Preference Model is a AI company based in San Francisco, USA. For a AI company this size, the reality is opportunity to shape your role based on the company stage.

Most Preference Model jobs are based in San Francisco, USA.

How to actually get hired at Preference Model

Why applying the normal way doesn't work

Preference Model uses an ATS, but referrals still come first. Cold applications aren't ignored β€” they're just behind referrals, sourced candidates, and recruiter picks.

Who to contact at Preference Model

Who decides
the hiring manager
Best channel
LinkedIn or direct email

What to show them

Build a small Karotte environment (uvx karotte create-env) for a kernel task where the grader checks only the outcome: match an fp64 reference within tolerance, as fast as possible, any precision allowed. Then document two hacks you tried against your own grader (for example, precompiled output pasted in, or timing manipulation) and how you closed them.

A cold email that works at Preference Model

Subject: Member of Technical Staff, ML Capabilities (build RL environments that teach models to do frontier-lab ML research/engineering, SF, $200K–$350K), [your one-line proof]
Hi Jennifer, your fp8-rmsnorm-gemm example stuck with me: the ban on Triton couldn't be graded, so the honest fix is to stop asking for it. I built a Karotte env for [kernel] that grades only against an fp64 reference, and here are the two hacks I closed on my own grader. Happy to walk through it.
Get personalised templates β†’

What Preference Model screens for

a16z leading a $16M seed for a company that has spent a year quietly selling to frontier labs is a bet that RL environments become a real supply chain, not a side project inside each lab. The angel list says the same: Ian Goodfellow, Julian Schrittwieser (AlphaGo/MuZero), Fei-Fei Li and Dylan Patel all have direct views into how frontier models get trained. The Kernels team posting is the most specific signal. They say low-level domains like GPU kernels, vector ISAs and FPGA work are where frontier models are weakest and underrepresented in training data, so anyone with real kernel experience is rare for them. They just went public and opened every role at once, so this is the window.

Tailor your CV to the specific Preference Model role rather than sending a general one. Applications that mirror the language of the job description clear automated filters at a materially higher rate.

avoid

Don't make these mistakes

Generic 'I'm passionate about AI safety' framing. Jennifer built data infrastructure, tokenizers and datasets on Anthropic's data team and has just published a detailed set of environment-design principles; vague alignment enthusiasm reads as someone who hasn't read them. Also skip 'I've built evals'. Their whole pitch is that evals and training environments fail differently. Lead with a specific reward hack you found and closed.

Mistakes that kill Preference Model applications

  • Don't send the same CV you sent everywhere else. At people it's obvious, and it's the fastest rejection there is.

  • Lead with their problem, not your ambition. Preference Model is focused on RL environments are now where frontier capability and alignment get decided, and the failure mode is concrete β€” show you understand that.

  • Don't apply and wait. The median AI application gets no response ever. One follow-up at day five roughly doubles reply rates.

tracking Preference Model

Applying to Preference Model? Get the contact, not the form.

notify me when Preference Model hires

The Preference Model interview process

4 stages14 days typicaltake-home: yesmodelled from similar companies

We don't yet have verified candidate reports for Preference Model. What follows is the typical process for a -person AI company, so treat it as a model, not confirmed detail.

Interview stages

  1. 1

    Recruiter Screen

    Phone or video Β· 30 min

    What it testsBasic qualification and logistics
    Usually run byRecruiter or HR
  2. 2

    Hiring Manager Interview

    Video call Β· 45 min

    What it testsRole fit and experience deep-dive
    Usually run byHiring manager
  3. 3

    Technical/Functional Round

    Video call Β· 60 min

    What it testsSkills assessment and problem-solving
    Usually run byTeam members
  4. 4

    Final Round

    In-person or video Β· 60 min

    What it testsCulture fit and cross-functional alignment
    Usually run bySenior leadership

Preference Model take-home assignment

Preference Model includes a take-home exercise in their interview process. For AI roles, this typically involves a practical problem that takes 2-4 hours. Focus on clean, working code over premature optimization. They're evaluating how you think and communicate, not just the solution.

Preference Model interview timeline

Preference Model runs about 14 days from first contact to offer. The median for AI companies at people is 14 days, so Preference Model is about average than most.

Interviewed at Preference Model?

Tell us how it went: stages, questions, timeline. Takes 90 seconds and it's how this page stays accurate for the next person.

Submit your Preference Model interview experience β†’

Preference Model jobs, frequently asked questions

How many jobs does Preference Model have open?

We're tracking 5 active openings at Preference Model (verified October 9, 2026).

Does Preference Model hire remotely?

All current Preference Model roles are based in San Francisco, USA.

What roles is Preference Model hiring for?

Preference Model is hiring across Engineering, Other, Data. The most recent opening is Member of Technical Staff, ML Capabilities (build RL environments that teach models to do frontier-lab ML research/engineering, SF, $200K–$350K).

How do I apply for a job at Preference Model?

Click through to apply, or see our detailed guide on landing a job at Preference Model.

Does Preference Model respond to cold emails?

We haven't verified response rates at Preference Model yet.

Who is the hiring manager at Preference Model?

At this size, hiring is usually run by the hiring manager.

How competitive is it to get hired at Preference Model?

Expect 100-250 applicants in the first two weeks for AI roles at this size. Apply within 72 hours for best odds.

How many rounds is the Preference Model interview?

4 stages: Recruiter Screen, Hiring Manager Interview, Technical/Functional Round, Final Round.

Is the Preference Model interview hard?

The interview emphasizes technical depth and system design over abstract problems. Hardest stage: Technical Interview.

Does Preference Model give a take-home task?

Yes, Preference Model includes a take-home assignment.

How long does Preference Model take to get back to you?

Around 14 days across the full process.

What should I prepare for the Preference Model interview?

Study technical depth and system design. At this size (), they care about self-sufficiency over textbook knowledge.

Where is Preference Model based?

Preference Model is headquartered in San Francisco, USA.

tracking Preference Modellast checked October 9, 2026

Get Preference Model roles before they're posted

A role stays uncontested for about four days. Here's the window, and where we put you in it.

notify me when Preference Model hires

From $5.99/week, cancel any time.

  1. livewatching Preference Model0 applicants
  2. +2hrole spotted & verified1
  3. +4hyou get the alert1
  4. day 3you've applied~8
  5. day 7+hits the job boards250+

Related

More AI companies hiring

Other startups we track in the same space.