Cookbook

Compiler presets shaped after public writeups — not quotes or endorsements. Each card maps onto this compiler: a verifier, a task, and a base model. Open it in Build to download the loop.

Glean · Tool plan, then cite

Waldo

DPO warmup on preferred vs. rejected tool calls, then an environment score for cited documents and thrift. Nemotron-3-Nano for latency.

DPO on production traces, then this RL loop.01 Verifiable outcome · Importance Sampling02 Specialized Tool Agents03 Nemotron-3-Nano Open in builder
Chroma · Search subagent

Context-1

Self-distill quality-filtered search trajectories, then reward retrieval plus token efficiency. A subagent that prunes its own context.

SDFT on teacher traces, then environment RL for recall.01 Demonstrations · Cross-Entropy (SDFT)02 Specialized Tool Agents03 Qwen3-8B Open in builder
Axiom Math · Kernel is the judge

AxiomProver

Formal math. Lean checks every step. The reward is binary: the proof verifies or it does not. Large base model, small adapter.

01 Verifiable outcome · Importance Sampling02 Formal Reasoning Engines03 DeepSeek-V3.1 Open in builder
Lightning Rod Labs · Calibration is the product

Clinical forecast

A lightweight adapter on a large base model for out-of-sample clinical outcomes. Reward the probability, not the essay.

01 Preference pairs · Pairwise preference02 Calibrated Forecasting03 DeepSeek-V3.1 Open in builder
Mantic · Decorrelated from the frontier

Judgmental forecast

RL on real-world forecasting questions. The fine-tune is useful in an ensemble because it is not a clone of GPT.

01 Verifiable outcome · Importance Sampling02 Calibrated Forecasting03 Qwen3.6-35B-A3B Open in builder
Shenfeld et al. · Learn without forgetting

SDFT skill add

The model is its own teacher. On-policy rewrite of a demonstration, then cross-entropy. New tools without wrecking the generalist.

01 Demonstrations · Cross-Entropy (SDFT)02 Specialized Tool Agents03 Nemotron-3-Nano Open in builder
Cookbook — Reinforcement: Build Your Own Reward Model