Evaluating Agentic AI Contribution
Incredible look at how Hamel's team evaluated Devin last year. Their research/data-rooted approach is awesome and gives a blueprint for how to eval agentic contribution.
https://www.answer.ai/posts/2025-01-08-devin.html
Crucial takeaways:
- They approached the work with a plan to evaluate performance
- They established task categories, and what success looked like (greenfield, research/analysis, code modification — plus bug fixes and refactors)
- They had a team working on this at the same time, spreading and sharing the knowledge
- They compared notes at the end and shared the failure modes uncovered