
2 days ago
TMax: Closing the Frontier Gap With Open Data
In this episode of Token Engineering, Cooper sits down with Yash, Head of AI Research at Neurometric AI, to break down TMax from AI2 (Allen Institute for AI)—a fully open dataset and training recipe that pushes small open-weight models like Qwen to near-frontier performance on terminal agent tasks. The paper argues that diversity and difficulty matter more than where the data comes from—even when that data is entirely invented by Gemini rather than pulled from real-world code.
We talked about:
- What TMax is, and why AI2 built a fully open recipe instead of a closed benchmark
- Why Gemini-invented data beat real GitHub repos on diversity
- Same recipe, opposite results: boosting one Qwen version, degrading the next
- How small models “cheat” when a task is beyond them
- Why a lighter harness outperformed a more complex one on Claude Haiku
- Gains that spilled past Terminal-Bench into SWE-bench and AIME math
- The compute cost of agentic RL training, and how runaway tool calls blow the budget
- Neurometric’s “Harbor Master,” built to run evals past AI2’s narrow benchmark set
Resources Mentioned:
TMax: A recipe for terminal agents: https://arxiv.org/abs/2606.23321
Connect with Neurometric:
Website: https://www.neurometric.ai/
Substack: https://neurometric.substack.com/
X: https://x.com/neurometricai/
Bluesky: https://bsky.app/profile/neurometric.bsky.social
Host/s:
Calvin Cooper
https://www.linkedin.com/in/coopernyc
Guest/s:
Yash Sharma
No comments yet. Be the first to say something!