FIG · 01OutcomeFinalpass / failProcessS1+.2S2+.8S3-.4S4+.6step scores isolate the failed reasoning stepProcess reward model
Post-training6 min

Process reward modeling: how step-level signals beat outcome-only RLHF

Aditi SharmaApril 18, 2026
FIG · 02Sandboxmocked UIBrowserlive tracegapmissing auth / latencykept tools / statereal browser traces expose the sandbox gapSandbox fidelity gap
RL environments9 min

Why your browser agent can't generalize: the sandbox-fidelity gap

Marcus ChenApril 11, 2026
FIG · 03High agreementShared missrubric gapagreement can hide a shared rubric missAgreement blind spot
Data quality7 min

Inter-annotator agreement is a lossy metric. Here's what we use instead.

Priya IyerApril 4, 2026
FIG · 04human labelsjudge scoreResidualsbias leftresidual bars expose judge driftJudge calibration
Evaluation5 min

LLM-as-judge calibration: a practitioner's checklist

David MwangiMarch 27, 2026
FIG · 05AgentpolicyEnvironmentapp stateactionrewardlive traces replace static rowsInteractive RL gym
Industry4 min

The shift from static datasets to interactive RL gyms

Vraify ResearchMarch 21, 2026
FIG · 06ChosenARejected BDPO oksingle turnUse RLHFstateful rewardstateful rewards need RLHFDPO / RLHF boundary
Post-training8 min

DPO vs RLHF: when chosen/rejected pairs are enough

Aditi SharmaMarch 14, 2026
Subscribe

Get research drops in your inbox.

Monthly. Substantive. Unsubscribe anytime.