
research note
Calibration Without Comprehension — Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software
This paper investigates whether fine-tuned large language models (LLMs) genuinely learn to reason about software vulnerabilities or simply calibrate their outputs without true comprehension










