Abstract
Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning. Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood. In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models. PHRBench characterizes each reasoning trajectory independently of final-answer correctness through Hallucination Compliance, Hallucination Avoidance, and Heuristic Correction, and defines an insightful trajectory as successful correction that ultimately reaches the correct answer. Across 4,820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory. We further find that properties of the hallucinated prompt contain substantial predictive signal for successful recovery, with a lightweight predictor achieving an AUROC of 0.847. These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.

Results across Models and Domains

Hallucinated context reduces accuracy by 7.7 percentage points on average. Larger models can exhibit both greater vulnerability to erroneous context and more frequent corrective behavior.
Table 1: Outcome performance and reasoning trajectory statistics under the truthful augmentation (T) and hallucinated augmentation (H) settings.
| Model | Acc (%) | OL | UI | BUF | BC | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| T | H | T | H | T | H | T | H | T | H | |
| Proprietary Models | ||||||||||
| Claude-Sonnet-4 | 57.5 | 48.4↓9.1 | 100.7 | 113.2 | 0.054 | 0.071 | 0.152 | 0.329 | 1.133 | 1.111 |
| GPT-4o-mini | 55.5 | 48.3↓7.2 | 71.2 | 72.4 | 0.072 | 0.059 | 0.130 | 0.186 | 0.752 | 0.720 |
| GPT-4o | 51.8 | 42.3↓9.5 | 60.0 | 60.8 | 0.061 | 0.053 | 0.113 | 0.214 | 0.659 | 0.618 |
| GPT-5.2 | 61.6 | 44.7↓16.9 | 43.0 | 45.1 | 0.007 | 0.016 | 0.013 | 0.026 | 0.578 | 0.515 |
| Gemini-2.0-Flash | 56.2 | 45.1↓11.1 | 76.2 | 86.5 | 0.046 | 0.061 | 0.092 | 0.178 | 0.745 | 0.785 |
| Gemini-2.5-Pro | 58.5 | 46.4↓12.1 | 82.4 | 95.3 | 0.045 | 0.068 | 0.124 | 0.244 | 0.852 | 0.850 |
| Open-source Models | ||||||||||
| Qwen2.5-1.5B-Instruct | 24.6 | 20.6↓4.0 | 115.6 | 117.1 | 0.272 | 0.218 | 0.091 | 0.076 | 1.987 | 1.748 |
| Qwen2.5-7B-Instruct | 44.5 | 38.0↓6.5 | 153.2 | 147.1 | 0.319 | 0.151 | 0.148 | 0.121 | 2.265 | 1.939 |
| Qwen2.5-32B-Instruct | 46.8 | 39.6↓7.2 | 228.4 | 215.1 | 0.245 | 0.150 | 0.285 | 0.353 | 2.820 | 2.644 |
| Qwen2.5-72B-Instruct | 52.2 | 44.2↓8.0 | 276.9 | 262.8 | 0.188 | 0.112 | 0.414 | 0.485 | 3.426 | 3.150 |
| Qwen3-8B | 46.0 | 43.8↓2.2 | 225.8 | 256.6 | 0.217 | 0.072 | 0.576 | 0.392 | 2.855 | 3.470 |
| Qwen3-32B | 48.5 | 41.0↓7.5 | 296.3 | 288.4 | 0.181 | 0.065 | 0.605 | 0.465 | 3.232 | 3.820 |
| Qwen3-235B-A22B | 58.2 | 47.4↓10.8 | 365.0 | 415.2 | 0.085 | 0.045 | 0.843 | 0.753 | 4.155 | 4.828 |
| Llama-3.1-8B-Instruct | 34.2 | 31.9↓2.3 | 210.5 | 199.9 | 0.378 | 0.322 | 0.371 | 0.207 | 2.752 | 2.222 |
| Llama-3.1-70B-Instruct | 48.4 | 40.9↓7.5 | 285.2 | 275.4 | 0.431 | 0.419 | 0.523 | 0.384 | 3.356 | 2.989 |
| Llama-3.2-3B-Instruct | 31.2 | 29.5↓1.7 | 167.0 | 153.4 | 0.159 | 0.137 | 0.322 | 0.124 | 2.076 | 1.441 |
| GLM-4-9B-Chat | 32.3 | 22.4↓9.9 | 127.3 | 79.0 | 0.171 | 0.173 | 0.050 | 0.038 | 1.886 | 1.122 |
| Mistral-7B-Instruct | 21.1 | 15.7↓5.4 | 131.0 | 122.4 | 0.130 | 0.104 | 0.262 | 0.370 | 1.853 | 1.472 |
T: truthful augmentation; H: hallucinated augmentation. Red annotations indicate the decrease in accuracy (percentage points). OL: Output Length; UI: Uncertainty Index; BUF: Belief Update Frequency; BC: Branching Complexity.
Insightful Trajectory Rates

Overview
- PHRBench contains 460 base questions, 460 truthful augmentations, and 3,900 hallucinated augmentations, yielding 4,820 controlled instances.
- It spans Chemistry, Biomedicine, Physics, and Code Generation, with base questions drawn from eight established benchmarks.
- Scientific questions have ten hallucinated variants each; coding problems have three. Each augmentation introduces a targeted semantic error while preserving the task.

Hallucinated Augmentation Types
| Type | Description |
|---|---|
| Rule Contradiction | A premise conflicts with an established rule, principle, or domain fact. |
| State Distortion | A statement is incorrect under the specific conditions of the current problem. |
| Pseudoscientific Entanglement | An outdated, discredited, or pseudoscientific theory is introduced as a premise. |
Behavioral Resolution
| Behavior | Description |
|---|---|
| Hallucination Compliance | The model accepts the hallucinated premise and reasons from it. |
| Hallucination Avoidance | The model bypasses the premise without explicitly correcting it. |
| Heuristic Correction | The model identifies and corrects the hallucinated premise. |
An insightful trajectory performs Heuristic Correction and reaches the correct final answer. Behavior is classified independently of final correctness.
Evaluation Metrics
The paper reports final-answer accuracy alongside Output Length (OL), Uncertainty Index (UI), Belief Update Frequency (BUF), and Branching Complexity (BC). Each model is sampled ten times per prompt at temperature 0.7. See Section 4.2 and Appendix B for definitions and configurations.
Qualitative Examples

Predicting Successful Recovery
Prompt-level structural and semantic features provide predictive signal before generation. The lightweight predictor achieves an AUROC of 0.847. SHAP attributions characterize predictive associations, rather than causal effects.
Global Feature Contributions

BibTeX
Citation with the author list provided by the research team; publication details will be updated with the paper release.
@misc{phrbench,
title = {PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs},
author = {Linghao Meng and Feng He and Xuan Yang and Junyuan Mao and Pinze Ren and Deqing Mu and Hesen Yang and Qiankun Li},
note = {Manuscript under review}
}