# ResearchMath-14K：通过智能体扩展研究级数学

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-05-27 08:00
- AIHOT 分数：70
- AIHOT 标记：精选
- AIHOT 链接：https://aihot.news/items/cmpovigyt09i3slv4yyi85evn
- 原文链接：https://arxiv.org/abs/2605.28003

## 精选理由

这可能是目前数学推理方向最有价值的数据集之一，它暴露了模型编造引用的问题，过滤后微调还能涨点，做数学推理的团队应该立刻拉下来试试。

## AI 摘要

本文介绍了ResearchMath-14K，这是一个包含14,056个研究级数学问题的数据集，通过多智能体流程从学术资料中策划而成，是目前此类规模最大的集合。研究还生成了ResearchMath-Reasoning（包含220K条教师轨迹），发现语言模型存在回避行为，且新一代模型产生的引用和虚假引用分别是旧模型的5.6倍和5.0倍。经过智能体过滤后，对参数规模为4B到30B的Qwen3模型进行微调，其平均得分比基础模型提高了9.2分，表明过滤后的开放问题尝试能为研究级数学推理提供有效监督。该数据集已公开发布。

## 正文

Guijin Son

Seungyeop Yi

Minju Gwak

Hyunwoo Ko

Wongi Jang

Youngjae Yu

Seoul National University

OneLineAI

Yonsei University

guijin.son@snu.ac.kr youngjaeyu@snu.ac.kr

Abstract

The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully engage with such problems without human intervention. A major obstacle is the lack of large-scale research-level math datasets. To this end, we introduce ResearchMath-14k, a set of problems curated from academic sources via a multi-agent pipeline, making it the largest collection of research-level mathematical problems to date. We further generate ResearchMath-Reasoning, K teacher trajectories from two open models, where we observe recurring avoidance behaviors such as non-attempts and fabricated references. Interestingly, across eight open-weight models, newer generations produce more references and more fake references per trace. After agentic filtering of ResearchMath-Reasoning, fine-tuning Qwen3 models from 4B to 30B parameters improves over base models by points on average. This shows that filtered open-problem attempts can provide useful supervision even without fully correct reasoning traces. We make ResearchMath-14k publicly available for future works on research-level mathemtical reasoning.111https://huggingface.co/datasets/amphora/ResearchMath-14k

ResearchMath-14k: Scaling Research-Level Mathematics via Agents

Guijin Son1,2 Seungyeop Yi1 Minju Gwak3 Hyunwoo Ko2 Wongi Jang1 Youngjae Yu1††thanks: Corresponding author Seoul National University1 OneLineAI2 Yonsei University3 guijin.son@snu.ac.kr youngjaeyu@snu.ac.kr

1 Introduction

Figure 1: Domain distribution of the ResearchMath-14k. Bubble size is proportional to the number of problems in each mathematical area. Logic and Foundations ( problems, ) and Other/Cross-disciplinary ( problems, ) are omitted from the visualization for readability.

Mathematicians are trained over years, escalating from undergraduate textbooks and exercises to seminar problems, qualifying-style questions, and short-term research. Over time, they learn practices that are central to becoming mathematicians: decomposing problems into lemmas, testing examples, isolating tractable subproblems, distinguishing a plausible route from a proof, and reasoning under genuine uncertainty. Frontier proprietary models increasingly appear to internalize parts of this curriculum (Alexeev et al., 2026a, b; Zheng et al., 2026). However, the open-source landscape has not kept pace. Nearly all publicly available math training data targets contest-style problems at the olympiad level or below (Li et al., 2024; Fan et al., 2025b), and the few datasets that do reach the research frontier are positioned as held-out evaluation benchmarks, often gate-kept to prevent contamination (Glazer et al., 2024; Phan et al., 2025).

Figure 2: Agentic construction pipeline for ResearchMath-14k. Starting from curated open-problem lists and research papers, an extractor agent maps document sections, detects candidate open questions, preserves supporting quotes, and rewrites them as standalone questions. A refiner agent then verifies open status, assigns taxonomy labels, and rewrites each candidate into a self-contained research problem before producing the final JSON record with statement, status, domain metadata, source, and solution fields when applicable.

Where, then, can research-level mathematical questions be obtained at scale? Recent work has largely relied on two expensive sources: multi-LLM pipelines that synthesize difficult problems (Zhang et al., 2026; Dekoninck et al., 2026), or expert mathematicians who write and curate them by hand (Son et al., 2026b; Garre et al., 2026). Both approaches are valuable, but neither provides an easy path to a broad, open training corpus. We take a different route. The mathematical literature already contains thousands of open problems, conjectures, seminar questions, and research directions. The bottleneck is extracting them from their local context and rewriting them into self-contained form. We collect 1,233 open-problem lists and research papers from zbMATH, arXiv, and academic repositories, then leverage agents to identify candidate questions, recover missing definitions and assumptions, and normalize them into standalone research-level problems. This process yields ResearchMath-14k, a corpus of 14,056 research-level mathematical questions along with ResearchMath-Reasoning, K reasoning trajectories generated from two open models.

In a manual review of 100 sampled trajectories, roughly 30% are visibly problematic, including non-attempts, substitutions to narrower problems, and fabricated arXiv or PDF URLs (Section 2). These failures recur in a larger trace-level analysis of eight open-weight models, including DeepSeek V4-Pro (DeepSeek-AI, 2026) and Kimi K2.6 (Team et al., 2026) (Section 3.2). Interestingly, newer models become more citation-heavy but less factual, with of 720 ResearchMath-14k traces containing at least one fake reference (Section 4). We use the same behavioral and factuality filters to clean ResearchMath-Reasoning into ResearchMath-Reasoning-Filtered, a -trace training-ready subset (Section 5). Fine-tuning three Qwen3 base models (4B, 8B, and 30B-A3B) on ResearchMath-Reasoning-Filtered improves them by percentage points on average, showing that the ResearchMath family is a valuable training resource for research-level reasoning even without ground-truth solutions. We openly release the ResearchMath family, comprising ResearchMath-14k ( research-level problems), and ResearchMath-Reasoning (K teacher trajectories), under the MIT license to support future work on research-level mathematical reasoning.

2 ResearchMath-14k

2.1 Collecting Existing Open Questions

We build ResearchMath-14k with a two-stage agentic pipeline: an Extractor agent pulls candidate problem statements from each source document, and a Refiner agent rewrites each statement into a self-contained problem, consulting online references. The pipeline produces problems from source documents, see Figure 2.

Sources.

Mathematicians have long published unresolved questions through workshops, surveys, and curated lists (Guy, 2004), both to attract collaborators and to record which questions a field regards as important enough to foreground for the broader community. Our pipeline captures both classical entries, such as Hilbert- or Erdős-style problem lists, and contemporary problem statements about modern mathematical objects and local technical settings. The latter are closer to the day-to-day research questions a working mathematician might pose at a workshop or in a recent survey, and are therefore the kind of supervision signal we target. See Appendix A for examples.

Specifically, source documents are drawn from three streams. arXiv open-problem papers (e.g., Diethelm et al. (2022)) ( documents, problems) are surveyed by searching arXiv for titles and abstracts mentioning “open problems,” or “unsolved.” Open-problem web pages ( documents, problems) are discovered by Google search and cover hosts such as academia.edu, MathOverflow, and Wikipedia. Problem-session sheets and curated lists ( documents, problems) are the third stream and include two sub-types: AIM-style workshop problem sessions where participants pose questions at the end of a meeting,222https://aimath.org/pastworkshops/nonselfadjointproblems.pdf and conference/proceedings open-problem rounds compiled by an editor at the close of a special session.

Extractor Agent.

The Extractor, driven by Codex with GPT-5.5 at xhigh reasoning effort, processes one source per run. It first follows the source URL down to the PDF or HTML page that holds the full text, discarding any document hidden behind a paywall. Before extraction, it also screens the document to confirm that it actually contains a problem list, skipping papers that do not in fact pose open problems (e.g., regular research papers that merely mention “open problem”). It then reads the paper end-to-end and extracts each open problem as a verbatim quote together with a first-level rewrite. While rewriting, the model is instructed to jump back and forth through the paper to pull in every definition and statement needed to understand that problem. Across the documents the Extractor yields a mean of questions per source (median , maximum ).

Refiner Agent.

Reading through the extracted questions, the authors noticed that some of the snippets still miss the definitions, notation, and hypotheses the original paper treats as already given. The Refiner, driven by Claude Code with Opus 4.7 at medium reasoning effort, fills that gap. It performs two tasks. First, it re-reads the original paper to inline every definition and hypothesis needed to state the problem in isolation. Second, it searches up to ten later papers that cite or extend the source, both to pull in the background of the source treated as implicit and to determine whether the problem has since been resolved. Each problem is tagged as open, partially solved, solved, or unknown. We audit random records labels using GPT-5.5 as LLM-Judge. It labels of refined statements as self-contained, compared with of original extractions, a percentage-point improvement (Appendix B). Refined statements also average characters, up from at the Extractor stage, a expansion.

媒体内容 · 前往原文查看

Dataset #Problems Source Diff.

GSM8K (Cobbe et al., 2021) k textbook

MATH (Hendrycks et al., 2021) k competitions

LeanDojo (Yang et al., 2023) k research lit.

MathInstruct (Yue et al., 2024) k synthetic

MetaMathQA (Yu et al., 2024) k synthetic

PRM800K (Lightman et al., 2024) k competitions

NuminaMath (Li et al., 2024) k synthetic

AceMath-Instruct (Liu et al., 2025b) M synthetic

OpenMathInstruct (Toshniwal et al., 2024) M synthetic

Riemann-Bench (Garre et al., 2026) expert

FrontierMath (Glazer et al., 2024) expert

Soohak (Son et al., 2026a) expert

GHOSTS (Frieder et al., 2023) textbook

HARDMath (Fan et al., 2025a) textbook

HLE (Phan et al., 2025) expert

ResearchMath-14k research lit.

Table 1: Representative public math datasets. Top section: training datasets are large, but do not expand to the research level. Lower section: existing research-grade resources are small (k items) and evaluation-only. ResearchMath-14k fills both gaps, large-scale and research-grade. Difficulty key: grade-school, olympiad, undergrad, research.

2.2 Filtering Near Duplicates

The collection pipeline produces a seed set of problems, but multiple sources often state the same open problem in slightly different forms, making duplicate filtering necessary. We embed all problems with Qwen3-Embedding-8B (Zhang et al., 2025) and compute pairwise similarities over both the original statements and the self-contained rewrites. Questions extracted from the same paper often share extensive background text and can look similar even when they are distinct, so a low similarity threshold would introduce many false positives. After manually inspecting borderline pairs at several cutoffs, we set the threshold to . A pair is marked as a duplicate if either similarity score exceeds this value. This threshold separates most true duplicates from same-paper false positives. For each duplicate pair, we keep the version from arXiv or another paper source and discard the version discovered through Google search; when both sources have the same priority, we choose one at random. Although this filtering is conservative, some distinct but closely related questions may still be removed, so we also release the raw seed set. This leaves a final collection of problems. See Appendix D for further details on the similarity distribution and examples of non-duplicate question pairs near the threshold.

2.3 Dataset Statistics

Composition.

Each problem is assigned a three-level taxonomy. The level-one domain groups are:

Each problem is also assigned one of macro-subjects and a research-level category tag ( unique tags). The hierarchy runs from broad area to research field to local topic. For example, one branch is:

Figure 1 shows the level-one distribution. The corpus is broad but skewed toward four large areas: Analysis/PDEs/Dynamics, Mathematical Physics, Discrete Mathematics/Combinatorics, and Geometry/Topology together account for problems (). A small fraction, problems (), falls into the Other/Cross-disciplinary group and covers science-adjacent open questions (e.g., on supernova progenitors, origin of language, computational theory of mind). Open problems form the majority (, ), followed by unknown (, ), partially solved (, ), and solved (, ). The set is source-diverse, spanning unique documents, with the top contributing problems () and the top contributing ().

Difficulty.

Difficulty is multidimensional, and a problem can be hard because it requires obscure background knowledge (Knowledge), demands novel thinking that deviates from existing approaches (Novelty), or involves compute-heavy multi-step reasoning (Procedural). We compare ResearchMath-14k against AceMath (Liu et al., 2025b), AIME(2024--2026) (Dekoninck et al., 2026), HLE-Verified (Phan et al., 2025), and NuminaMath (Li et al., 2024). From each of the five datasets we sample problems and consider all dataset pairs. For each pair we randomly draw cross-dataset problem pairs and randomize their order, giving total comparisons. Each comparison is judged by GPT-5-mini along the three axes, producing win/loss/draw labels from which we compute Elo ratings. On all three axes, ResearchMath-14k ranks above these existing math datasets by roughly Elo points (Figure 3), implying that it is a qualitatively harder problem class rather than an incremental step above existing math datasets. This highlights our contribution as the hardest open-source math problem set to date.

Figure 3: Elo Ratings for Difficulty Comparison. Ratings are computed from pairwise LLM difficulty judgments. All sources start at with ; wins score , losses score , and draws score . Higher Elo means the source is judged more difficult more often.

2.4 Generating Responses

We use two teacher models, GPT-OSS-120B (Agarwal et al., 2025) and Qwen3-30B-A3B (Yang et al., 2025), to generate reasoning trajectories for ResearchMath-14k. Note that the goal is not to produce correct solutions. Most solutions are not yet known, and we do not expect sub-trillion-parameter models to solve open research questions. We initially fine-tune Qwen3-4B on these trajectories without any filtering. This leads to substantial degeneration of the student model, including repetitive outputs and frequent non-attempts.333We do not report specific scores for this unfiltered fine-tune because the resulting model degenerated on nearly every evaluation, scoring close to zero. The point of the anecdote is the failure mode, which motivates the larger-scale analysis. To understand why, we conduct a human review of randomly sampled trajectories. We find that in cases the teacher does not attempt the problem at all. Instead, the model appears to recognize the question as an open problem and outputs a non-attempt in one of the following forms:

: lists known related references, and outputs “open” as the answer.

: after concluding the problem is open, narrows the conditions and either solves the narrowed version or simply lists related references.

These observations motivate the larger-scale behavioral and factuality analysis in Section 3.2. Nonetheless, the resulting set pairs K prompts with K responses (approximately per prompt) from two teacher models, and we release it as ResearchMath-Reasoning, which is, to our knowledge, the largest publicly available collection of model attempts on research-level math.

3 Experiment Setup

The cause of such fabricated reasoning trajectories (Section 2.4) is subject to several possible explanations. The behavior may reflect problem difficulty, stylistic mismatch between paper-derived prompts and benchmark-style questions, or the limited capacity of GPT-OSS-120B. We therefore set up experiments across models and benchmarks (Section 3.1) and evaluate them with complementary behavioral metrics (Section 3.2).

3.1 Baselines

Models.

We evaluate a broad set of models, including several substantially larger systems and both older and newer generations from each model family: DeepSeek R1 (Guo et al., 2025), DeepSeek V4-Pro (DeepSeek-AI, 2026), Kimi K2 (Team et al., 2025), Kimi K2.6 (Team et al., 2026), Qwen3 (30B-A3B, 235B-A22B) (Yang et al., 2025), and Qwen3.5 (35B-A3B, 397B-A17B) (Qwen Team, 2026). Throughout the analysis we group these models into four oldernewer matched pairs (R1V4-Pro, K2K2.6, Qwen3 30BQwen3.5 35B, and Qwen3 235BQwen3.5 397B).

Benchmarks.

ResearchMath-14k has two defining properties: problems are research-level, and their surface form is AI-refined from a source paper. We choose four control benchmarks to isolate each property. To control for any artifact of the AI-refining step, we use SOOHAK (Son et al., 2026a) and Leipzig Tier-4 (ScienceBench, 2026), both research-level but human-authored. To study the effect of difficulty, we use the math subset of HLE-Verified (a version of Humanity’s Last Exam (Phan et al., 2025) verified by Zhai et al. (2026)) and AIME (Zhang and Math-AI, 2024, 2025, 2026). Both are easier than the research-level sets, with AIME being easiest. AIME combines questions from 2024, 2025, and 2026 for problems in total. We sample items from each of the other four benchmarks, with SOOHAK restricted to items labeled graduate or beyond from the challenge subset, and all benchmarks further filtered to short-form-answer questions; this leaves SOOHAK with items, for prompts overall.

3.2 Behavior and Factuality Metrics

Analyzing trace-level behavior is not trivial. We use two complementary methods that together cover two aspects of a reasoning trace, the model’s behavior (how it reasons) and the factuality of what it cites. Each method covers both axes.

Rule-Based Counting.

We use three curated phrase lists, each targeting a distinct phrasing pattern. The lists were assembled by the authors after reviewing dozens of model reasoning traces and collecting recurrent phrases that fit each pattern, and matching is performed against the lowercased trace (full lists in Appendix C). cite matches citation-like nouns (e.g. “paper”). abandon catches abandonment (e.g. “cannot solve”, “educated guess”). assume catches claims made without justification (e.g. “known result”, “i remember”). Two of these (abandon, assume) measure behavior, while cite measures factuality and bridges into the agent-judge below. Each counter increments by one per match, and per benchmark we report the row-hit rate , the fraction of traces in which counter matches at least once ( is the match count in trace and is the set of traces). These rules are transparent, cheap, and chosen to broadly cover recurring failure patterns. Counting alone, however, cannot judge whether a given match is a real failure in context.

Figure 4: Citation behavior across four matched oldernewer model pairs. Left: Row-hit-rate deltas (newer minus older, in percentage points) for the three rule-based counters (abandon, assume, cite) across the five benchmarks. Right: Agent-Judge reference verification on 720 ResearchMath-14k traces, one point per model. Newer models (orange) sit upper-right of their predecessors (blue), with more reference-like mentions per trace (x-axis) and more references judged fake per trace (y-axis); dashed guides mark 10% and 20% fake-reference shares. Full per-model counts in Appendix E.

Agent-Judge.

For an additional behavior check, we use GPT-5.5 as a judge (Zheng et al., 2023) to detect lemma decomposition. The judge is prompted to generate a binary label on whether the solver model breaks the problem into provable subgoals, inspected over the first of the trace, where subgoal-setting tends to happen. We highlight lemma decomposition as it is one of the most critical behaviors for LLMs to tackle open questions across long reasoning time. The factuality check inspects whether reference-like spans in the trace correspond to real sources. Because running an agent over a full reasoning trace is expensive, we use a two-stage pipeline. We slice each trace into newline-delimited blocks and use GPT-5.4-nano as to audit each block and extract reference-like spans (books, papers, website URLs). A search-enabled Codex agent then iterates over each span to confirm whether the span is genuine reference text (filtering out e.g. named mathematical theorems) and whether the referenced source exists on the web. We provide the surrounding block for reference, and require multiple web searches before every judgment. Prompts for both checks are in Appendix H.4. Both judge outputs measure properties of the reasoning trace, not correctness.

4 Analyzing Reasoning Behavior on ResearchMath-14k

The manual review in Section 2 flagged roughly of teacher trajectories as visibly problematic. We now measure the same failure modes at corpus scale using the eight models and five benchmarks from Section 3, and report two findings.

Citation-like reasoning rises sharply in newer model generations (Figure 4, left, cite row), with row-hit rates increasing by 30-80 percentage points on ResearchMath-14k, Leipzig Tier-4, and SOOHAK across the DeepSeek, Kimi, and Qwen3 matched pairs. The effect weakens as benchmarks get easier (modest on HLE, near zero on AIME), suggesting that newer models’ tendency to cite is an artifact of the academic level of the questions.

To supplement the keyword counter, we use the Agent-Judge (Section 3.2) on 90 traces from ResearchMath-14k for each of our 8 models. Across all 720 traces, 629 (87.4%) cite at least one reference-like object and 389 (54.0%) contain at least one fake reference. At the reference level, we inspect 19,864 extracted mentions and label 3,492 fake (17.6%) after consulting internet search (Figure 4, right). Per-trace mention counts grow dramatically across the matched comparisons. DeepSeek R1 V4-Pro rises from to mentions per trace ( fakes), Kimi K2 K2.6 from to ( fakes), Qwen3 30B Qwen3.5 35B from to ( fakes), and Qwen3 235B Qwen3.5 397B from to ( fakes). In aggregate, newer models produce more reference-like mentions per trace and more fakes. The fake mentions are mostly hallucinated paper titles and author attributions. Models try to ground their arguments on wrong statements by fabricating that a supporting reference exists, making the result sound correct. Representative fakes:

“Neeman’s paper: A remark on the unique factorization theorem”

“J. Winkelmann, On the holomorphic equivalence of the Koras–Russell cubic”

“a specific paper: On the probability that a random polynomial is stable by J. M. Anderson”

Why do newer models fabricate more often?

Interestingly, we observe that models released in 2025 (DeepSeek R1, Kimi K2, Qwen3) cite less, while models released in 2026 (DeepSeek V4-Pro, Kimi K2.6, Qwen3.5) cite far more, with more fake citations. In other words, factuality on research-level prompts is moving backward. Because this pattern holds across DeepSeek, Kimi, and Qwen, three different model families, it is unlikely to be a quirk of any single training set. One plausible explanation is internet-search RL, or more broadly agentic RL (Dong et al., 2025; Liu et al., 2025a; Li et al., 2026). Recent post-training pipelines often place the model inside an agentic harness at train time, equipped with explicit search and citation tools, and reward it for grounding claims in retrieved sources. Over training, the model learns to invoke papers, books, and URLs as a routine part of producing an authoritative-looking answer. In our setting, however, models are evaluated without internet access. A plausible explanation is that rather than abandoning the citation behavior when the search tool is unavailable, models keep invoking the learned pattern and simply fabricate the references they would normally retrieve.

It should be noted, however, that citations and compression are not themselves failures. Mathematicians cite, reduce, and skip routine details too, and if models could ground their citations correctly, this would be less of a concern. But citations are not the only place where models try to look the part. On ResearchMath-14k the abandon counter (Section 3.2) matches only traces (), while assume matches (; Figure 4, left). Models rarely give up outright, and the attempt almost always leans on compressed claims rather than from-scratch derivation. These surface signs resemble mathematician practice, but we cannot tell whether models employ the underlying reasoning or simply parrot the form.

媒体内容 · 前往原文查看

Model ResearchMath-14k Leipzig T4 SOOHAK

DeepSeek R1

DeepSeek V4-Pro

Kimi K2

Kimi K2.6

Qwen3 30B

Qwen3.5 35B

Qwen3 235B

Qwen3.5 397B

Table 2: LLM-judge lemma-decomposition positives by model and benchmark. Each cell reports positive traces over judged traces; the colored parenthetical on the newer model gives the delta (newer minus older) within each matched pair (green = newer fires more, red = newer fires less).

Using the Agent-Judge lemma-decomposition metric from Section 3.2, we find that the behavior is almost absent (Table 2). Across ResearchMath-14k, Leipzig Tier-4, and SOOHAK, only judged traces are marked positive, and on ResearchMath-14k only . This matters not just for ResearchMath-14k but for any research-level mathematics models will face. Such problems are hard enough that they cannot be solved in a single pass and must be broken down into checkable subproblems.

5 Learning from ResearchMath-14k

Prior work shows that supervised fine-tuning on mathematical reasoning can tolerate a moderate fraction of incorrect solutions (Toshniwal et al., 2025; Muennighoff et al., 2025; Son et al., 2025). We push this idea to a setting where correctness is largely unavailable. For ResearchMath-14k, most problems are open or beyond the reach of current sub-trillion-parameter LLMs, so the trajectories in ResearchMath-Reasoning are unlikely to be complete nor correct. However, we hypothesize that, after filtering ResearchMath-Reasoning to exclude trajectories flagged by either rule-based counters or agent judges (Section 3.2), the remaining traces are qualitatively different from merely wrong work. Watching a trained researcher attempt an open problem and fall short can be instructive in a way that watching a kindergarten student make an arithmetic mistake is not. The former may introduce relevant objects, explore plausible reductions, test examples, or develop partial arguments, while the latter usually carries little transferable structure. Whether these wrong-but-reasonable traces are genuinely useful is therefore an empirical question with practical consequences. Requiring verified-correct reasoning at the research level would mean expert-annotating every trajectory, a cost that does not scale. If such traces are sufficient to teach useful behavior, they provide a cheaper path for future frontier-level data curation. In the following section, we investigate whether training on these attempts provides a useful signal.

5.1 Training Setup

We filter ResearchMath-Reasoning using the Agent-Judge pipeline from Section 3.2. This verifies every reference-like span against web search, and traces containing any reference judged fake are removed. Because the agent step calls multiple agents and paid web-search APIs, our budget allows producing only filtered traces, which form ResearchMath-Reasoning-Filtered. For comparison, we randomly sample traces from DASD-Thinking (Yan et al., 2026) to test the alternative explanation that any gain comes from learning the output format rather than from research-level content. We fine-tune Qwen3-4B/8B/30B-A3B-base with LoRA on each training set. See Appendix G for training configurations. We evaluate on AIME 2024–2026 (), HLE (), and SOOHAK Challenge and Mini combined (). We filter HLE and SOOHAK to include questions with integers only, and use math-verify444https://github.com/huggingface/Math-Verify for scoring.

5.2 Training Results

Figure 5: Fine-tuning results by benchmark. Bars show mean score for each model averaged across three runs; whiskers show standard deviation over three runs.

Training on ResearchMath-Reasoning-Filtered improves over the base models in all modelbenchmark cells, with a mean gain of percentage points, while DASD improves in . ResearchMath-Reasoning-Filtered also outperforms DASD in cells, with the clearest gains on the research level evaluations. Averaged over HLE and SOOHAK, it is points above DASD, with the largest gaps at HLE for the 30B model and SOOHAK for the 4B model . The only exception is AIME for the 30B model, where DASD wins by points.

Two implications follow. First, the gains are not explained by generic math reasoning exposure alone. DASD improves the base models, but ResearchMath-Reasoning-Filtered does better in nearly all settings. Second, useful research-level supervision need not be verified-correct. Once non-attempts, unsupported claims, and fake citations are removed, wrong-but-reasonable attempts still improve student models. We leave further tests of this signal at larger scale to future work.

6 Related Works

Research-Level Mathematics with LLMs.

Inducing mathematical reasoning in LLMs has been driven mainly by resources with known answers (Toshniwal et al., 2024; Li et al., 2024; Yuan et al., 2026). However, most remain below the research frontier. These problems are typically solved, verifiable (Albalak et al., 2025), synthetic (Toshniwal et al., 2025), textbook-derived (Fan et al., 2025b), olympiad-derived (Mahdavi et al., 2025; Ko et al., 2025), or tied to formal proof environments (Yang et al., 2023). Research-level mathematical data remains expensive and nontrivial to scale: existing resources are often expert-authored (Son et al., 2026a), private or gated (Garre et al., 2026), small, continuously maintained for evaluation (Dekoninck et al., 2026), or difficult to convert into training material because usable prompts require local definitions, notation, hypotheses, status checks, and deduplication (Zhang et al., 2026). We address this gap by collecting research-level questions already present in the mathematical literature and rewriting them. The result is ResearchMath-14k, a -problem corpus that, to the best of our knowledge, is the largest collection of research-level mathematical questions available for training.

7 Conclusion and Future Work

This work uses ResearchMath-14k to study how open reasoning models behave on research-level mathematical problems whose complete solutions are often unavailable. Our trace-level analysis shows a concerning shift. Newer model generations produce more citation-heavy responses, but also more fake references. At the same time, these imperfect attempts still contain useful supervision. Fine-tuning on the filtered trajectories improves models by an average percentage points over their base versions. These results suggest that research-level training need not rely only on verified complete solutions: wrong-but-reasonable attempts can be useful when their most harmful failure modes are removed. We encourage future works to test this signal at a larger scale, while clarifying when correct traces remain necessary for reliable proof behavior. We publicly release ResearchMath-14k and ResearchMath-Reasoning to support future works on research-level mathematics.

References

S. Agarwal, L. Ahmad, J. Ai, S. Altman, A. Applebaum, E. Arbus, R. K. Arora, Y. Bai, B. Baker, H. Bao, et al. (2025) Gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925. Cited by: §2.4.

A. Albalak, D. Phung, N. Lile, R. Rafailov, K. Gandhi, L. Castricato, A. Singh, C. Blagden, V. Xiang, D. Mahan, et al. (2025) Big-math: a large-scale, high-quality math dataset for reinforcement learning in language models. arXiv preprint arXiv:2502.17387. Cited by: §6.

B. Alexeev, M. Putterman, M. Sawhney, M. Sellke, and G. Valiant (2026a) Short proofs in combinatorics and number theory. arXiv preprint arXiv:2603.29961. Cited by: §1.

B. Alexeev, M. Putterman, M. Sawhney, M. Sellke, and G. Valiant (2026b) Short proofs in combinatorics, probability and number theory ii. arXiv preprint arXiv:2604.06609. Cited by: §1.

K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. (2021) Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. Cited by: Table 1.

DeepSeek-AI (2026) DeepSeek-v4: towards highly efficient million-token context intelligence. Cited by: §1, §3.1.

J. Dekoninck, N. Jovanović, T. Gehrunger, K. Rögnvaldsson, I. Petrov, C. Sun, and M. Vechev (2026) Beyond benchmarks: matharena as an evaluation platform for mathematics with llms. External Links: 2605.00674, Link Cited by: §1, §2.3, §6.

K. Diethelm, V. Kiryakova, Y. Luchko, J. T. Machado, and V. E. Tarasov (2022) Trends, directions for further research, and some open problems of fractional calculus. Nonlinear Dynamics 107 (4), pp. 3245–3270. Cited by: §2.1.

G. Dong, H. Mao, K. Ma, L. Bao, Y. Chen, Z. Wang, Z. Chen, J. Du, H. Wang, F. Zhang, et al. (2025) Agentic reinforced policy optimization. arXiv preprint arXiv:2507.19849. Cited by: §4.

F. Fan, S. Martinson, E. Wang, K. Hausknecht, J. Brenner, D. Liu, N. Peng, C. Wang, and M. Brenner (2025a) Hardmath: a benchmark dataset for challenging problems in applied mathematics. In International Conference on Learning Representations, Vol. 2025, pp. 13523–13556. Cited by: Table 1.

R. Fan, Z. Wang, and P. Liu (2025b) Megascience: pushing the frontiers of post-training datasets for science reasoning. arXiv preprint arXiv:2507.16812. Cited by: §1, §6.

S. Frieder, L. Pinchetti, C. Chevalier, R. Griffiths, T. Salvatori, T. Lukasiewicz, P. Petersen, and J. Berner (2023) Mathematical capabilities of chatgpt. Advances in neural information processing systems 36, pp. 27699–27744. Cited by: Table 1.

S. Garre, E. Knutsen, S. Mehta, and E. Chen (2026) Riemann-bench: a benchmark for moonshot mathematics. arXiv preprint arXiv:2604.06802. Cited by: §1, Table 1, §6.

E. Glazer, E. Erdil, T. Besiroglu, D. Chicharro, E. Chen, A. Gunning, C. F. Olsson, J. Denain, A. Ho, E. d. O. Santos, et al. (2024) Frontiermath: a benchmark for evaluating advanced mathematical reasoning in ai. arXiv preprint arXiv:2411.04872. Cited by: §1, Table 1.

D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Ding, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Chen, J. Yuan, J. Tu, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. You, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Zhou, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang (2025) DeepSeek-r1 incentivizes reasoning in llms through reinforcement learning. Nature 645 (8081), pp. 633–638. External Links: Document, ISBN 1476-4687, Link Cited by: §3.1.

R. K. Guy (2004) Unsolved problems in number theory. Vol. 21, Springer. Cited by: §2.1.

D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt (2021) Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874. Cited by: Table 1.

H. Ko, G. Son, and D. Choi (2025) Understand, solve and translate: bridging the multilingual mathematical reasoning gap. In Proceedings of the 5th Workshop on Multilingual Representation Learning (MRL 2025), pp. 78–95. Cited by: §6.

J. Li, E. Beeching, L. Tunstall, B. Lipkin, R. Soletskyi, S. Huang, K. Rasul, L. Yu, A. Q. Jiang, Z. Shen, et al. (2024) Numinamath: the largest public dataset in ai4maths with 860k pairs of competition math problems and solutions. Hugging Face repository 13 (9), pp. 9. Cited by: §1, §2.3, Table 1, §6.

W. Li, B. Qu, B. Pan, J. Zhang, Z. Liu, P. Zhang, W. Chen, and B. Zhang (2026) LiteResearcher: a scalable agentic rl training framework for deep research agent. arXiv preprint arXiv:2604.17931. Cited by: §4.

H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe (2024) Let’s verify step by step. In International Conference on Learning Representations, Vol. 2024, pp. 39578–39601. Cited by: Table 1.

J. Liu, Y. Li, C. Zhang, J. Li, A. Chen, K. Ji, W. Cheng, Z. Wu, C. Du, Q. Xu, et al. (2025a) Webexplorer: explore and evolve for training long-horizon web agents. arXiv preprint arXiv:2509.06501. Cited by: §4.

Z. Liu, Y. Chen, M. Shoeybi, B. Catanzaro, and W. Ping (2025b) Acemath: advancing frontier math reasoning with post-training and reward modeling. In Findings of the Association for Computational Linguistics: ACL 2025, pp. 3993–4015. Cited by: §2.3, Table 1.

S. Mahdavi, M. Li, K. Liu, C. Thrampoulidis, L. Sigal, and R. Liao (2025) Leveraging online olympiad-level math problems for llms training and contamination-resistant evaluation. arXiv preprint arXiv:2501.14275. Cited by: §6.

N. Muennighoff, Z. Yang, W. Shi, X. L. Li, L. Fei-Fei, H. Hajishirzi, L. Zettlemoyer, P. Liang, E. Candès, and T. B. Hashimoto (2025) S1: simple test-time scaling. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 20286–20332. Cited by: §5.

L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, et al. (2025) Humanity’s last exam. arXiv preprint arXiv:2501.14249. Cited by: §1, §2.3, Table 1, §3.1.

Qwen Team (2026) Qwen3.5: towards native multimodal agents. External Links: Link Cited by: §3.1.

ScienceBench (2026) Benchmarks in leipzig. Note: Accessed: 2026-05-22 External Links: Link Cited by: §3.1.

G. Son, S. Kim, C. Arnett, H. Ko, H. Lee, H. Kang, J. Longxi, J. Yun, J. Lee, K. Lee, et al. (2026a) Soohak: a mathematician-curated benchmark for evaluating research-level math capabilities of llms. arXiv preprint arXiv:2605.09063. Cited by: Table 1, §3.1, §6.

G. Son, D. Yang, H. L. Patel, A. Agarwal, H. Ko, C. Lim, S. Panda, M. Kim, N. Drolia, D. Choi, et al. (2025) Pushing on multilingual reasoning models with language-mixed chain-of-thought. arXiv preprint arXiv:2510.04230. Cited by: §5.

G. Son, D. Yang, H. L. Patel, H. Ko, A. Agarwal, S. Ahn, K. Lee, and Y. Yu (2026b) Judging what we cannot solve: a consequence-based approach for oracle-free evaluation of research-level math. arXiv preprint arXiv:2602.06291. Cited by: §1.

K. Team, T. Bai, Y. Bai, Y. Bao, S. Cai, Y. Cao, Y. Charles, H. Che, C. Chen, G. Chen, et al. (2026) Kimi k2. 5: visual agentic intelligence. arXiv preprint arXiv:2602.02276. Cited by: §1, §3.1.

K. Team, Y. Bai, Y. Bao, Y. Charles, C. Chen, G. Chen, H. Chen, H. Chen, J. Chen, N. Chen, et al. (2025) Kimi k2: open agentic intelligence. arXiv preprint arXiv:2507.20534. Cited by: §3.1.

S. Toshniwal, W. Du, I. Moshkov, B. Kisacanin, A. Ayrapetyan, and I. Gitman (2025) Openmathinstruct-2: accelerating ai for math with massive open-source instruction data. In International Conference on Learning Representations, Vol. 2025, pp. 19243–19275. Cited by: §5, §6.

S. Toshniwal, I. Moshkov, S. Narenthiran, D. Gitman, F. Jia, and I. Gitman (2024) Openmathinstruct-1: a 1.8 million math instruction tuning dataset. Advances in Neural Information Processing Systems 37, pp. 34737–34774. Cited by: Table 1, §6.

S. Yan, K. Liu, C. Shen, B. Wang, S. Fan, J. Zhang, Y. Wu, Z. Wang, and J. Ye (2026) Distribution-aligned sequence distillation for superior long-cot reasoning. arXiv preprint arXiv:2601.09088. Cited by: §5.1.

A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al. (2025) Qwen3 technical report. arXiv preprint arXiv:2505.09388. Cited by: §2.4, §3.1.

K. Yang, A. Swope, A. Gu, R. Chalamala, P. Song, S. Yu, S. Godil, R. J. Prenger, and A. Anandkumar (2023) Leandojo: theorem proving with retrieval-augmented language models. Advances in Neural Information Processing Systems 36, pp. 21573–21612. Cited by: Table 1, §6.

L. Yu, W. Jiang, H. Shi, J. Yu, Z. Liu, Y. Zhang, J. Kwok, Z. Li, A. Weller, and W. Liu (2024) Metamath: bootstrap your own mathematical questions for large language models. In International Conference on Learning Representations, Vol. 2024, pp. 45040–45061. Cited by: Table 1.

W. Yuan, J. Yu, S. Jiang, K. Padthe, Y. Li, D. Wang, I. Kulikov, K. Cho, Y. Tian, J. Weston, et al. (2026) Naturalreasoning: reasoning in the wild with 2.8 m challenging questions. Advances in Neural Information Processing Systems 38. Cited by: §6.

X. Yue, X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen (2024) Mammoth: building math generalist models through hybrid instruction tuning. In International Conference on Learning Representations, Vol. 2024, pp. 40320–40341. Cited by: Table 1.

W. Zhai, Z. Wang, J. Wang, B. Yang, X. Li, X. Xu, B. Wang, P. Wang, X. Wu, A. Li, et al. (2026) HLE-verified: a systematic verification and structured revision of humanity’s last exam. arXiv preprint arXiv:2602.13964. Cited by: §3.1.

J. Zhang, C. Petrui, K. Nikolić, and F. Tramèr (2026) Realmath: a continuous benchmark for evaluating language models on research-level mathematics. Advances in Neural Information Processing Systems 38. Cited by: §1, §6.

Y. Zhang, M. Li, D. Long, X. Zhang, H. Lin, B. Yang, P. Xie, A. Yang, D. Liu, J. Lin, et al. (2025) Qwen3 embedding: advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176. Cited by: §2.2.

Y. Zhang and T. Math-AI (2024) American invitational mathematics examination (aime) 2024. Cited by: §3.1.

Y. Zhang and T. Math-AI (2025) American invitational mathematics examination (aime) 2025. Cited by: §3.1.

Y. Zhang and T. Math-AI (2026) American invitational mathematics examination (aime) 2026. Cited by: §3.1.

D. Zheng, I. von Glehn, Y. Zwols, I. Beloshapka, L. Buesing, D. M. Roy, M. Wattenberg, B. Georgiev, T. Schmidt, A. Cowie, et al. (2026) AI co-mathematician: accelerating mathematicians with agentic ai. arXiv preprint arXiv:2605.06651. Cited by: §1.

L. Zheng, W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al. (2023) Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems 36, pp. 46595–46623. Cited by: §3.2.

Appendix Contents

Appendix A Example Source Comparisons

Tables 3 and 4 illustrate the source-level contrast discussed in Section 2.1. The first pair stays within number theory/arithmetic geometry; the second pairs a broad algebraic-geometry grand challenge with a narrower modern question about hyperkähler Chow rings.

媒体内容 · 前往原文查看

Grand-challenge entry Contemporary workshop-style entry

Dataset record google__mppc/q_001 07-workshop-problems/q_007

Source The Millennium Prize Problems Some Open Problems About Diophantine Equations

Area Number theory; elliptic curves and -functions Number theory; rational points on hyperelliptic curves

Problem form Prove the Birch–Swinnerton-Dyer conjecture: for an elliptic curve , the order of vanishing of at equals the Mordell–Weil rank of . Prove unconditionally that the genus- curve y^2=-3x^6-x^5+2x^4+2x^2-3x-3 has no rational points over , given that this is known assuming BSD for its Jacobian.

Status Open Unknown

Comment This is the kind of classical, high-visibility challenge problem that appears in many public lists of unsolved mathematics. This is narrower and more local: it sits near BSD and rational-points methods, but asks for a concrete unconditional proof for one curve. It better represents the workshop and survey questions that form much of ResearchMath-14k.

Source Clay monograph PDF Leiden workshop PDF

Table 3: Side-by-side example of a classical arithmetic-geometry grand-challenge source and a narrower contemporary open-problem source represented in ResearchMath-14k.

媒体内容 · 前往原文查看

Grand-challenge entry Contemporary survey-style entry

Dataset record google__mppc/q_004 1002.4321/q_010

Source The Millennium Prize Problems arXiv problem list on compact hyperkähler manifolds

Area Algebraic geometry; Hodge theory and algebraic cycles Algebraic geometry; hyperkähler manifolds and Chow rings

Problem form Prove the Hodge conjecture: every rational Hodge class on a smooth projective complex variety is a -linear combination of cohomology classes of algebraic subvarieties. For every projective hyperkähler manifold , prove that the subring of generated by and the Chern classes of injects into cohomology under the cycle-class map.

Status Open Partially solved

Comment This is a universal conjecture about the relation between Hodge-theoretic and algebraic cycles across all smooth projective complex varieties. This is much more local: it asks for a specific Chow-ring injectivity statement in the modern setting of hyperkähler geometry. It uses current objects and tools while still sitting near the Hodge/cycle-theoretic grand challenge.

Source Clay monograph PDF arXiv:1002.4321

Table 4: Second side-by-side example: a broad algebraic-geometry grand challenge compared with a narrower contemporary hyperkähler/Chow-ring problem represented in ResearchMath-14k.

Appendix B Self-Containment Audit

To quantify whether refinement makes questions usable without the source document, we run a first-pass automatic audit on randomly sampled released records. For each record, Codex labels both the original extracted question and the refined standalone question as self-contained or not. A statement is counted as self-contained only if a mathematically trained reader can understand the task from the text alone, without source-local notation, missing definitions, or references to external sections, figures, or problem numbers. Table 5 shows that original extracted questions are self-contained in of sampled cases, while refined standalone questions are self-contained in . The refiner turns initially non-self-contained snippets into self-contained questions, leaving refined questions flagged for remaining context gaps.

媒体内容 · 前往原文查看

Audit item Count Rate

Original extracted question

Refined standalone question

Original no, refined yes

Original yes, refined no

Table 5: Automatic self-containment audit. We sample released records and ask Codex to label both the original extracted question and the refined standalone question as self-contained or not.

媒体内容 · 前往原文查看

Case Original extraction judged not self-contained Refined standalone problem

Geodesics Suppose as . Let be a complete, simply connected Riemannian manifold without conjugate points. Let be two distinct unit-speed geodesics, and denote the Riemannian distance by . Suppose that as . Does it follow that as ?

Words Let be an alphabet and let be a -free word. Then is a factor of a maximal -free word over . Let be a finite alphabet and let denote the set of finite words over . A pattern is a non-empty word over an alphabet of variables. An occurrence of in a word is a factor of of the form , where is a non-erasing morphism. A word is -free if no factor of is an occurrence of . A -free word is maximal if it cannot be extended on the left or right while remaining -free. Prove or disprove: for every alphabet , every pattern , and every -free word , there exists a maximal -free word over that contains as a factor.

Remaining gap Fix with . Is it true that, with high probability for , the width of the independence complex satisfies ? Let be a finite simple graph with independence complex , whose faces are the independent sets of . Following Meshulam’s recursive framework, one associates to a nonnegative integer parameter and a finite sequence produced by the standard recursive process used to bound the topological connectivity of . For fixed and with , is it true with high probability that ? The refined version still depends on the source for the exact recursive definition of and .

Table 6: Examples from the self-containment audit. The first two rows show originally non-self-contained extractions that become self-contained after refinement. The final row shows a remaining failure case in which both the original extraction and refined statement still depend on source-local definitions.

Appendix C Keyword and Judge Metric Details

This appendix specifies the surface-form counters used in Section 3.2. All keyword matching is performed after lowercasing the analyzed text. A keyword group count is the sum of exact substring occurrences of all phrases in that group. The counters are descriptive trace features; they are not used as a standalone hallucination classifier.

C.1 Keyword Groups

abandon.

This counter marks claims of being stuck, time-limited, or unable to complete the solution. The keyword list is: lack of progress, given the time, time constraints, too complex, not practical manually, i’m stuck, i am stuck, dead end, can’t solve, cannot solve, without progress, hazard a guess, educated guess.

cite.

This counter marks references to source objects, external databases, or citation-like artifacts. The keyword list is: paper, book, article, textbook, monograph, survey, journal, proceedings, publication, arxiv, doi, wikipedia, mathworld, oeis, stackexchange, aops, art of problem solving, website, webpage, online source.

assume.

This counter combines assertive shortcuts and remembered-result language, both of which substitute confident assertion for derivation in the trace. The keyword list is: it can be shown, one can show, it is easy to see, clearly, obviously, intuitively, by symmetry, must be, should be, known result, standard result, i remember, similar problem online, look it up mentally, the problem implies, strongly suggests, well-known, it is known, i recall.

C.2 LLM-Judge and Agent-Judge Annotations

Agent-Judge reference verification.

This annotation is applied to traces that mention citation-like references, including papers, books, articles, arXiv identifiers, DOI-like strings, named sources, or database references. A Codex-based search agent queries the internet for the mentioned reference and records whether the referenced source appears to exist. The purpose is to separate genuine provenance signals from hallucinated bibliographic support. This annotation does not judge whether the source proves the model’s claim; it only checks whether the cited object itself can be found.

LLM-Judge lemma decomposition.

GPT-5.5 inspects the trace and marks whether the model decomposes the problem into explicit intermediate lemmas, claims, subgoals, or cases that structure the solution. A positive annotation requires more than generic planning language: the trace should state a reusable intermediate fact or subproblem and then use it in the subsequent reasoning. The purpose is to measure constructive proof organization rather than surface verbosity.

LLM-Judge counterexample search.

GPT-5.5 inspects the trace and marks whether the model actively tests a conjecture, proposed formula, candidate solution, or simplifying assumption against counterexamples, edge cases, small instances, or adversarial constructions. A positive annotation requires an explicit attempt to falsify or stress-test an idea, not merely checking arithmetic. The purpose is to measure whether the model uses skeptical reasoning before committing to a claim.

C.3 Aggregation

For each keyword group, LLM-Judge annotation, and Agent-Judge verification result, we aggregate by model, model family, and benchmark. The row-hit rate treats each trace as a binary hit for a counter; for judged annotations, this is the fraction of traces marked positive. The benchmark trend view reports the average newer-minus-older delta across the DeepSeek, Kimi, and Qwen comparison pairs.

媒体内容 · 前往原文查看

Figure 6: Pairwise Similarity Distribution. Empirical distributions of pairwise embedding similarities across all problem pairs. Dashed vertical lines indicate the maximum similarity values, both below the duplicate threshold. Left: similarities between original statements. Right: similarities between self-contained rewrites.

Appendix D Near-Duplicates Filtering Details

D.1 Pairwise Similarity Distribution

Figure 6 reports the pairwise embedding similarity distributions for the original statements and the self-contained rewrites in ResearchMath-14k.

D.2 GPT-5.5 Judgments Near Decision Boundary

Tables 7, 8, and 9 present problem pairs with similarity scores close to the threshold that GPT-5.5 nevertheless judged to be distinct. These cases support our conservative threshold choice.

Table 7:

GPT-5.5

Judgment Example 1

Table 8:

Judgment Example 2

Table 9:

GPT-5.5

Judgment Example 3

Lemma-decomposition rows

ResearchMath-14k

Leipzig Tier-4

SOOHAK

HLE-Verified

AIME

Research-level total

Table 10:

Absolute behavior-counter row-hit rates across the eight paper models. The

assume

cite

abandon

counters correspond to the three rule-based keyword groups defined in Appendix

C.1

. Lemma-decomposition rows are judged separately by the Agent-Judge and are available for the three research-level benchmarks.

Appendix E Reasoning Behavior Details

E.1 Behavior-Counter Rates

Table 10 reports the fraction of traces in each benchmark that trigger each behavior counter across the eight evaluated models.

Appendix F License and Release

We release the ResearchMath family under the MIT License. The release covers two artifacts: ResearchMath-14k, the corpus of research-level mathematical problems described in Section 2, and ResearchMath-Reasoning, the K reasoning trajectories described in Section 2.4. Both artifacts derive from publicly available academic sources (arXiv preprints, open-problem web pages, and workshop or conference problem sheets). The Extractor agent discards any document hidden behind a paywall before extraction (Section 2.1); paywalled or restricted sources are not represented in the released data.

Appendix G Training Details

All fine-tuning runs use LoRA on top of three Qwen3 base models (Qwen3-4B-base, Qwen3-8B-base, Qwen3-30B-A3B-base) on randomly sampled traces from either filtered ResearchMath-14k or the DASD-Thinking control. Each setting is run with three seeds and the reported numbers are averages over the runs.

LoRA configuration.

Rank , alpha , dropout , no bias, applied to the attention and MLP projections of each transformer block (q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj).

Batching.

Per-device batch size is ; global batch size is for the B run and for the smaller models.

Sequence length.

Examples are truncated at tokens for the B/B runs and tokens for the B run.

Appendix H Prompts

H.1 Dataset Generation Agents

Table 11 and Table 12 present the prompts used for the Extractor and Refiner agents in Section 2.1, respectively.

H.2 Difficulty Comparison

Table 13 shows the prompt for difficulty comparison in Section 2.3.

H.3 Response Generation

Table 14 shows the prompt used to generate model responses in Section 2.4.

H.4 Factuality Metrics

Tables 15, 16, and 17 present the prompts used for the factuality metric in Section 3.2. Table 17 gives the short-block variant of Table 16, which is used for detected blocks shorter than 200 characters with additional surrounding reasoning context.

Table 11:

Prompt for the Extractor Agent

Table 12:

Prompt for the Refiner Agent

Table 13:

Prompt for Difficulty Comparison

Table 14:

Prompt for Response Generation

Table 15:

Prompt for factuality reference-span extraction

Table 16:

Prompt for factuality agent verification

Table 17:

Short-block prompt for factuality agent verification
