Overmind 发布一个可将代码与 traces 转为 context graph、构建训练与评测数据并微调小型专用模型的工具,代码开源于 github.com/overmind-core/overmind。
A frontier model reading a contract will sometimes cite a clause that isn't there.
Overmind fine-tunes a small open model on your own production traces, then scores it against your current model on evals built from those same real tasks.
Their published comparison is Qwen3.5 9B tuned through Overmind against GPT5.6 Luna. The company reports 20 to 30x fewer phantom clauses in legal contracts and 7x better accuracy at quoting a clause word for word.
Here the observability layer and the training layer share the same data. The traces that show you where the agent fails become the dataset, and the evals are built from real tasks rather than a public benchmark.
The output is a smaller open model with weights you own. You can host it on Overmind or run it yourself.
Overmind turns anyone into an AI lab. • Builds a context graph from code and traces • Curates training and eval datasets • Evaluates prompts and models • Trains smaller, specialized models https://github.com/overmind-core/overmind在 X 查看被引用的帖子
来源:Rohan Paul · x.com