2 days ago

TMax: Closing the Frontier Gap With Open Data

In this episode of Token Engineering, Cooper sits down with Yash, Head of AI Research at Neurometric AI, to break down TMax from AI2 (Allen Institute for AI)—a fully open dataset and training recipe that pushes small open-weight models like Qwen to near-frontier performance on terminal agent tasks. The paper argues that diversity and difficulty matter more than where the data comes from—even when that data is entirely invented by Gemini rather than pulled from real-world code.



We talked about:

 

  • What TMax is, and why AI2 built a fully open recipe instead of a closed benchmark
  • Why Gemini-invented data beat real GitHub repos on diversity
  • Same recipe, opposite results: boosting one Qwen version, degrading the next
  • How small models “cheat” when a task is beyond them
  • Why a lighter harness outperformed a more complex one on Claude Haiku
  • Gains that spilled past Terminal-Bench into SWE-bench and AIME math
  • The compute cost of agentic RL training, and how runaway tool calls blow the budget
  • Neurometric’s “Harbor Master,” built to run evals past AI2’s narrow benchmark set



Resources Mentioned:

TMax: A recipe for terminal agents: https://arxiv.org/abs/2606.23321 




Connect with Neurometric:
Website: https://www.neurometric.ai/ 

Substack: https://neurometric.substack.com/ 

X: https://x.com/neurometricai/ 

Bluesky: https://bsky.app/profile/neurometric.bsky.social

 

Host/s:

Calvin Cooper

https://x.com/cooper_nyc_ 

https://www.linkedin.com/in/coopernyc

 

Guest/s:

Yash Sharma

https://x.com/yash_j_sharma 

https://www.linkedin.com/in/yashjsharma

Comment (0)

No comments yet. Be the first to say something!

Copyright 2025 All rights reserved.

Podcast Powered By Podbean

Version: 20241125