NVIDIA 在 Megatron Core 中降低比特级确定性预训练开销
Scale Bitwise-Deterministic Pretraining with NVIDIA Megatron Core
NVIDIA 在 Megatron Core 中降低比特级确定性预训练的开销,并以万亿参数 Nemotron 模型为案例验证。该方案提供独立可复现与比特级断点续训两项保证,通过逐步指纹比对训练损失、梯度范数等指标定位首个不一致步骤,并借助从迭代到 kernel 的渐进式追踪排查分歧来源。
Bitwise determinism makes large-scale pretraining easier to debug, validate, and resume reproducibly. These benefits become especially valuable when training models with trillions of parameters across thousands of GPUs, where multiple parallelism dimensions, low-precision computation, and distributed checkpointing complicate failure reproduction and fix validation.
Production and hero training runs are expensive and require a strong guarantee of success. Bitwise determinism enables reliable failure replay, loss-spike debugging, system-change validation, and recovery from interrupted training runs. At a trillion-parameter scale, even a small slowdown can cost thousands of GPU-days, so deterministic training must remain efficient. A trillion-parameter Nemotron case study shows how NVIDIA is reducing the overhead of bitwise-deterministic pretraining in NVIDIA Megatron Core.
What is bitwise determinism?
Bitwise determinism requires independent and checkpoint-resumed runs to follow the same numerical trajectory when data order, architecture, recipe, parallelism, software, runtime settings, and hardware are fixed. Megatron Core targets two guarantees.
Independent reproducibility: two runs launched from the same initial state must remain bitwise identical at every training step.
Bitwise checkpoint resume: a run that saves and restores one or more checkpoints must remain identical to an uninterrupted run.

Why should pretraining be deterministic?
Bitwise determinism provides four benefits for large-scale pretraining.
- Replay failures: Reproduce a loss spike at the same step, then change one component at a time to isolate the cause.
- Preserve checkpoint continuity: Resume interrupted training without changing the numerical trajectory.
- Detect unexpected differences: Use mismatches between otherwise identical runs to investigate checkpoint corruption or software and hardware issues.
- Compare changes: Distinguish the numerical effect of an intended change from run-to-run variation.

How to verify determinism
Printed loss values are not sensitive enough to establish bitwise identical numerics. Two runs can print the same rounded loss while differing in the lower-order bits of their gradients or parameters.
A stronger validation procedure fingerprints selected numerical metrics at every step. These can include training loss, language-modeling loss, load-balancing loss, multi-token prediction (MTP) loss, and gradient norm. Run the same configuration twice and compare these fingerprints step by step. The first mismatching step helps narrow the investigation.
To validate the bitwise checkpoint resume, compare an uninterrupted reference run with checkpoint-resumed runs. After every restore, verify that the parameters, optimizer state, RNG state, data position, and subsequent outputs are bitwise identical to the reference.
The trillion-parameter Nemotron model checkpoint path preserves a full-precision source of truth and reconstructs the runtime representation through the same conversion path used during training.


How to fix determinism when it breaks
At scale, nondeterminism may be intermittent or topology-dependent, and visible loss divergence can occur long after the first differing bit. The workflow proposed in Megatron-LM PR #7262 records ordered, per-rank tensor fingerprints for offline comparison. First, confirm identical seeds, data order, batch size, parallelism, container, and software stack.
Localize divergence with granular tracing
Begin with an end-to-end comparison, then narrow the capture scope from the training iteration to the phase, module, operation, and kernel. If two runs enter a scope with bitwise-identical inputs but leave it with different outputs, the divergence originated within that scope.
- End-to-end metrics detect the failure and identify the first iteration where a serialized metric differs.
- Broad semantic tracing covers collectives, pipeline communication, recomputation, optimizer operations, and gradient reductions to identify the divergent training phase.
- Module and layer tracing narrows the divergence to a particular model component, transformer layer, mixture of experts (MoE) block, or optimizer stage.
- Operation-level tracing fingerprints ATen operations and targeted extension calls to identify the first operation with matching inputs and different outputs.
- Kernel-level tools identify the responsible kernel, algorithm configuration, and nondeterministic mechanism.
Do not trace only the iteration where the loss visibly separates; the first differing bit may have appeared earlier. Use broad tracing to locate the earliest divergent iteration and ranks, then enable detailed operation and kernel tracing only within that narrowed scope.

Compare traces offline
Each selected rank writes its trace to a file without adding collectives or cross-rank ordering. This reduces the risk of masking the race being investigated. Align events by a run-independent identity, such as the operation name, occurrence count, and module scope.
Look for matching input fingerprints and different output fingerprints.
hash(input_run_A) == hash(input_run_B) hash(output_run_A) != hash(output_run_B)
“First mismatch” refers to each rank separately; sequence numbers cannot establish ordering across ranks. Classify each rank’s first mismatch by causal role.
| First mismatch on a rank | Interpretation | Next action |
|---|---|---|
| Inputs match; outputs differ | Candidate origin | Investigate this operation or its hidden producer |
| Inputs and outputs differ | Downstream receiver | Continue tracing upstream |
| No origin on traced ranks | Origin lies outside the capture | Widen the iteration window or rank set |
If many ranks identify the same originating operation, the operation itself is likely nondeterministic. If only a subset does, investigate topology, rank placement, input distribution, and reduction ordering.
Fingerprint tensors on the GPU
The proposed workflow uses torch.hash_tensor for GPU-resident fingerprints. Record each tensor’s shape, dtype, and element count alongside its digest. For MXFP8 or NVFP4 tensors, fingerprint both the encoded values and their scale buffers.
Whole-tensor XOR fingerprints cannot detect permutations. For order-sensitive data such as routing maps and MoE dispatch outputs, fingerprint these tensors by row or chunk using the dim argument. A fingerprint is an efficient check, but a match does not guarantee that the tensors are bitwise identical. Use byte-level comparisons to confirm bitwise equality.
Rule out false alarms
Check for tracing blind spots and invalid reads before identifying the cause.
- Dispatcher blind spots.
TorchDispatchModeobservesATenoperations routed through the PyTorch dispatcher. Custom kernels can escape this tracing scope. If the first mismatch appears at a simple view, slice, or addition, probe the custom kernel that produced its input. - Probe artifacts. Uninitialized tensors, incomplete asynchronous collectives, and nonblocking copies may be read before their contents are valid. Exclude these cases before declaring the producer nondeterministic.
Validate the fix
Validate the patch with paired-run checks.
- The original pair should reproduce the divergence.
- The patched pair should remain bitwise identical.
- Determine whether the fix corrects the nondeterministic implementation or routes execution around it.
- Measure the new performance cost.
Note that bitwise determinism is validated within the same hardware and software environment. Comparisons across GPU generations, network configurations, or library versions are outside this validation scope and may produce different results.
Optimizing deterministic training for a trillion-parameter Nemotron model
A deterministic recipe with substantial overhead may support debugging, but will be too costly for production training. Deterministic and nondeterministic execution should be co-optimized from the beginning. This work uses a trillion-parameter Nemotron model that combines Mamba-style state-space model (SSM) layers and Transformer attention layers. Deterministic training must cover SSM and attention kernels, low-precision computation, distributed communication, and transitions between layer types.
Establish a controlled baseline
Compare deterministic execution with the fastest supported nondeterministic recipe using the same model, hardware, batch sizes, parallelism strategy, precision format, software environment, and measurement window.
Measure:
- Throughput and step time
- Peak memory and GPU utilization
- Exposed communication time
- Time spent in major kernel groups
Calculate the determinism tax:
Determinism tax=(deterministic step timebaseline step time−1)×100%
Collect data only after training performance stabilizes.
Optimization journey

Optimization across three workloads reduced determinism overhead from double-digit baselines to low single digits. At 2,432 GPUs, the large-scale Nemotron recipe measured an approximately 2% steady-state determinism tax, with bitwise determinism over 800 steps.
Kernel-level optimization example
In the grouped-GEMM epilogue, multiple N-tiles originally accumulated into the same dprob[token] address, making the result depend on their arrival order. Serializing those writers restored determinism but reduced parallelism.
Give each N-tile a private output slot, preserve parallel execution inside the kernel, and combine the slots in a fixed order after all writers finish.

Validate optimization at scale
A low determinism tax on a few GPUs does not guarantee the same result at production scale. Communication, synchronization, pipeline bubbles, expert routing, and load balance can change the relative overhead. For a hypothetical run with a 100-day nondeterministic baseline on 10,000 GPUs, reducing overhead from 15% to 5% would save 10 days, or 100,000 GPU-days.
Maintain determinism as models, kernels, and recipes evolve
New kernels, fusions, precision formats, or parallelism configurations can introduce nondeterminism. Megatron-LM protects against regressions with these checks.
- Recipe validation:
--deterministic-modeapplies canonical environment settings, enables PyTorch deterministic algorithms, and rejects features without a deterministic path. - Kernel testing: Kernel tests repeat operations with identical inputs and RNG state, then compare outputs and gradients byte for byte.
- Module validation: Module-level tests repeat models and transformer blocks with restored RNG state across parallelism configurations, including FP8 and FP4. Independent end-to-end runs then verify that full-precision training metrics remain bitwise identical.
For current coverage and known gaps, see determinism status, operation catalog, and kernel testing guide.
Common sources of nondeterminism
Use Table 2 to identify likely causes of nondeterminism and choose the next diagnostic step.
| Observed symptom | Likely cause | Diagnostic method | Resolution |
|---|---|---|---|
| Runs diverge from the first step, but not consistently | Runtime autotuning selects different kernel configurations | Compare selected configurations and trace the earliest affected output | Pin or cache a validated configuration |
| MoE runs diverge only when a fused auxiliary loss is enabled | Reduction order is not fixed | Disable individual fusions and trace the loss computation | Use a deterministic reduction or disable the fusion |
| A resumed run gradually separates from the continuous run | Incomplete RNG state restoration | Compare RNG state immediately before and after resume | Save and restore every RNG stream bitwise |
| Divergence appears after a software or container update | A low-level library kernel changed | Bisect the stack and build a minimal kernel reproducer | Select a deterministic kernel path until corrected |
| Parameters differ immediately after loading a checkpoint | Low-precision weights or scales were reconstructed differently | Compare values and scaling metadata across save/load | Preserve a full-precision source of truth and exact reconstruction path |
| A recipe passes at small scale but fails at high scale | Scale-dependent communication or expert-parallel path | Increase scale systematically and trace selected ranks | Isolate and correct the first scale-dependent operation |
Get started
Bitwise determinism makes large-scale pretraining easier to debug and resume. The Nemotron case study shows how kernel and recipe optimizations can reduce its performance cost. Validate reproducibility and overhead in the hardware and software environment you plan to use.
To get started, review the Megatron Core User Guide for a supported recipe, then compare independent and checkpoint-resumed runs to validate reproducibility.
来源:NVIDIA Technical Blog:Agentic AI / Generative AI · developer.nvidia.com