Loading companies…
United States · AI
Reinforcement-learning environments that help AI labs train models to solve tasks without exploiting flaws in the grading system.
Preference Model develops training environments for frontier AI labs. Its work focuses on making the task, allowed actions and reward calculation agree, so training reinforces useful behaviour rather than shortcuts through a faulty evaluator.
Its open-source Karotte framework isolates model-generated code and separates execution from grading. It clears background processes and checks suspicious file types before scoring, addressing ways an agent might interfere with the evaluator instead of completing the assigned work.
| Announced | Round | Amount | Investors |
|---|---|---|---|
| Seed | $16M | Andreessen Horowitz (a16z) — LeadDr. Fei-Fei Li — ParticipantIan Goodfellow — ParticipantJulian Schrittwieser — ParticipantScale Angel Group — ParticipantSignalFire — ParticipantSouth Park Commons — Participant |