# Cohere CEO Aidan Gomez 称大模型公司用生产用户数据衍生合成数据训练

- 来源：Aidan Gomez (@aidangomez)
- 发布时间：2026-09-09 01:49
- AIHOT 分数：53
- AIHOT 链接：https://aihot.news/items/cmtszt7yc04m6ro5w45yx8hf0
- 原文链接：https://x.com/aidangomez/status/2097381789039837637

## AI 摘要

Cohere CEO Aidan Gomez 称从多个大实验室员工处听说，消费级 AI 工具的生产用户数据被转化为合成数据用于训练，处理复杂数学、商业、软件、生物问题等模型最需要学习的用例更易被上采样。

## 正文

Synthetic data derived from production user data of consumer AI tools is used for training. I’ve heard this rumour from both large labs’ employees.

In particular, if you’re doing something “interesting” like working on complex math/business/software/bio problems you’re dramatically more likely to get trained on because they filter/up-weight towards those usecases where the model has the most to learn.

Even in ZDR and “we won’t train on you” regimes, derivative data is usually carved out. The promise is only not to train on exactly the data you put in, rewritten data is fair game.

### 引用推文

> OpenAI：We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work th...
