Choose a verifier, a task, and a base model. Download a ready-to-run training loop.

Reinforcement: Build Your Own Reward Model

Input

You get

WaitingSelect a training signal. That chooses how success is scored.
Task and base model specialize the loop. The signal is not itself a trained reward model.
Build
0 / 3
Waldo

Tool agent

Reward cited tool calls and thrift.

Open recipe
AxiomProver

Reasoning model

Reward proofs that verify.

Open recipe
Judgmental forecast

Forecasting model

Reward calibrated probabilities.

Open recipe
Reinforcement: Build Your Own Reward Model