Training Data Reduction with High Fidelity Labels

The Gemini team published an interesting approach for enabling training of LLM evals (for fine-tuning or reinforcement learning) with smaller datasets. Using an LLM judge for initial open/axial coding phases and then HITL to further refine. They got a 100,000 data set to < 500.

This approach may be a cheat code for doing at least a small degree of RL on startup time/resource budget.

https://research.google/blog/achieving-10000x-training-data-reduction-with-high-fidelity-labels/