PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs

Linghao Meng1,*Feng He2,*Xuan Yang1,*Junyuan Mao1
Pinze Ren3Deqing Mu4Hesen Yang1Qiankun Li5,†
1National University of Singapore2Independent Researcher3Tsinghua University
4Johns Hopkins University5Nanyang Technological University

*Equal contribution.†Corresponding author.

Abstract

Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning. Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood. In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models. PHRBench characterizes each reasoning trajectory independently of final-answer correctness through Hallucination Compliance, Hallucination Avoidance, and Heuristic Correction, and defines an insightful trajectory as successful correction that ultimately reaches the correct answer. Across 4,820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory. We further find that properties of the hallucinated prompt contain substantial predictive signal for successful recovery, with a lightweight predictor achieving an AUROC of 0.847. These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.

PHRBench overview: hallucinations propagate through downstream systems; compliance, avoidance, and correction differ even when final outcomes agree; recovery shows scaling tensions and predictable signatures.
Figure 1. From hallucinated context to downstream recovery. PHRBench separates resolution behavior from final correctness. Click any figure to view at full resolution.

Experimental Results

Results across Models and Domains

Heatmap and radar plots of accuracy retention across 18 models, four domains and three hallucination types.
Accuracy retention. Performance varies with both the domain and hallucination type. Retention is accuracy under hallucinated context relative to truthful context, expressed as a percentage; it is not absolute accuracy.

Hallucinated context reduces accuracy by 7.7 percentage points on average. Larger models can exhibit both greater vulnerability to erroneous context and more frequent corrective behavior.

Table 1: Outcome performance and reasoning trajectory statistics under the truthful augmentation (T) and hallucinated augmentation (H) settings.

ModelAcc (%)OLUIBUFBC
THTHTHTHTH
Proprietary Models
Claude-Sonnet-457.548.4↓9.1100.7113.20.0540.0710.1520.3291.1331.111
GPT-4o-mini55.548.3↓7.271.272.40.0720.0590.1300.1860.7520.720
GPT-4o51.842.3↓9.560.060.80.0610.0530.1130.2140.6590.618
GPT-5.261.644.7↓16.943.045.10.0070.0160.0130.0260.5780.515
Gemini-2.0-Flash56.245.1↓11.176.286.50.0460.0610.0920.1780.7450.785
Gemini-2.5-Pro58.546.4↓12.182.495.30.0450.0680.1240.2440.8520.850
Open-source Models
Qwen2.5-1.5B-Instruct24.620.6↓4.0115.6117.10.2720.2180.0910.0761.9871.748
Qwen2.5-7B-Instruct44.538.0↓6.5153.2147.10.3190.1510.1480.1212.2651.939
Qwen2.5-32B-Instruct46.839.6↓7.2228.4215.10.2450.1500.2850.3532.8202.644
Qwen2.5-72B-Instruct52.244.2↓8.0276.9262.80.1880.1120.4140.4853.4263.150
Qwen3-8B46.043.8↓2.2225.8256.60.2170.0720.5760.3922.8553.470
Qwen3-32B48.541.0↓7.5296.3288.40.1810.0650.6050.4653.2323.820
Qwen3-235B-A22B58.247.4↓10.8365.0415.20.0850.0450.8430.7534.1554.828
Llama-3.1-8B-Instruct34.231.9↓2.3210.5199.90.3780.3220.3710.2072.7522.222
Llama-3.1-70B-Instruct48.440.9↓7.5285.2275.40.4310.4190.5230.3843.3562.989
Llama-3.2-3B-Instruct31.229.5↓1.7167.0153.40.1590.1370.3220.1242.0761.441
GLM-4-9B-Chat32.322.4↓9.9127.379.00.1710.1730.0500.0381.8861.122
Mistral-7B-Instruct21.115.7↓5.4131.0122.40.1300.1040.2620.3701.8531.472

T: truthful augmentation; H: hallucinated augmentation. Red annotations indicate the decrease in accuracy (percentage points). OL: Output Length; UI: Uncertainty Index; BUF: Belief Update Frequency; BC: Branching Complexity.

Insightful Trajectory Rates

Insightful trajectory rates: Qwen3-235B 0.25, Qwen3-32B 0.20, Qwen3-8B and Qwen2.5-72B 0.16; the lowest is GLM-4-9B at 0.01.
Insightful trajectory rates. Successful correction varies substantially across the 18 evaluated models. Values are rounded as in the supplied figure.

PHRBench Benchmark

Overview

460 base questions: 120 Chemistry, 120 Biomedicine, 120 Physics, and 100 Code Generation. Scientific hallucinations are predominantly rule contradictions; code hallucinations are state distortions.
Dataset composition. Base-question reasoning categories and hallucinated augmentation types. Chemistry, Biomedicine, and Physics each include 120 base questions; Code Generation includes 100.

Hallucinated Augmentation Types

TypeDescription
Rule ContradictionA premise conflicts with an established rule, principle, or domain fact.
State DistortionA statement is incorrect under the specific conditions of the current problem.
Pseudoscientific EntanglementAn outdated, discredited, or pseudoscientific theory is introduced as a premise.

Behavioral Resolution

BehaviorDescription
Hallucination ComplianceThe model accepts the hallucinated premise and reasons from it.
Hallucination AvoidanceThe model bypasses the premise without explicitly correcting it.
Heuristic CorrectionThe model identifies and corrects the hallucinated premise.

An insightful trajectory performs Heuristic Correction and reaches the correct final answer. Behavior is classified independently of final correctness.

Evaluation Metrics

The paper reports final-answer accuracy alongside Output Length (OL), Uncertainty Index (UI), Belief Update Frequency (BUF), and Branching Complexity (BC). Each model is sampled ten times per prompt at temperature 0.7. See Section 4.2 and Appendix B for definitions and configurations.

Qualitative Examples

Qualitative ion-identification example: a compliant trajectory predicts manganese, while a corrective trajectory updates its belief and correctly identifies cadmium.
Qualitative example. The supplied case contrasts compliance with a successful corrective trajectory. The belief update is the key behavioral distinction.

Predicting Successful Recovery

Prompt-level structural and semantic features provide predictive signal before generation. The lightweight predictor achieves an AUROC of 0.847. SHAP attributions characterize predictive associations, rather than causal effects.

Global Feature Contributions

SHAP feature contributions led by total prompt length at 20.7%, with candidate and question length at 8.8% each.
Global feature contributions. Total prompt length accounts for 20.7% of normalized feature attribution. These are predictive associations, not causal effects.

BibTeX

Citation with the author list provided by the research team; publication details will be updated with the paper release.

@misc{phrbench,
  title = {PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs},
  author = {Linghao Meng and Feng He and Xuan Yang and Junyuan Mao and Pinze Ren and Deqing Mu and Hesen Yang and Qiankun Li},
  note = {Manuscript under review}
}