About
Reinforcement.tech is a compiler for training loops. Choose a verifier, a task, and a base model. You download Tinker training code.
The first question is how you will score an output. A verifier returns a reward. That reward is the training signal — Lean, pytest, a preference pair, a demonstration you cannot afford to forget. It is not yet a learned reward model. The task and the base model specialize the loop.
Made by David Smooke via Cursor, with the newspaper-on-a-classic-Mac atmosphere of HackerNoon. Writing lives on the Reinforcement HackerNoon page. Icons are from the Pixel Icon Library.
Training runs on Tinker by Thinking Machines Lab. This site is not affiliated with, sponsored by, or endorsed by Thinking Machines Lab. Their mark is used to credit the API the compiled file calls.