ISSN :2582-9793

When the Audit Depends on the Auditor: Prompt Sensitivity in Matched-Pair Audits of Large Language Models

Original Research (Published On: 21-Jul-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.64322

JIahe Sui and Xiangzhi Li

Adv. Artif. Intell. Mach. Learn., - (-):-

1. JIahe Sui: City of London Freemen's School

2. Xiangzhi Li: Nanshan Administrative Institute, Shenzhen

Download PDF Here

DOI: 10.54364/AAIML.2026.64322

Article History: Received on: 10-Apr-26, Accepted on: 14-Jul-26, Published on: 21-Jul-26

Corresponding Author: JIahe Sui

Email: jiahesui91@gmail.com

Citation: Jiahe Sui and Xiangzhi Li. When the Audit Depends on the Auditor: Prompt Sensitivity in Matched-Pair Audits of Large Language Models. Advances in Artificial Intelligence and Machine Learning. 2026. (Ahead of Print) https://dx.doi.org/10.54364/AAIML.2026.64322


Abstract

    

Bias audits of large language models (LLMs) typically use a single prompt to compare model responses to identity-matched stimuli. This design rests on the implicit assumption that the estimated identity gap is stable across plausible prompt wordings. We test this assumption in a factorial experiment on three contemporary LLMs (Qwen, Llama, and GPT-4o), which evaluated ten matched-text managerial vignettes under four prompts crossed with gender and name-signalled ethnicity manipulations, yielding 48,000 model calls. Prompt choice moved both rating levels and, in some cells, the estimated identity gap. Prompt effects on rating levels were often larger than identity effects, especially when a deliberately critical stress-test prompt was included. Where the gap was non-trivial, sign reversals were rare; the substantive importance of the prompt-induced instability we observed depends strongly on whether the underlying gap is itself non-trivial. We summarise cross-prompt variation with the Audit Disagreement Index (ADI). Single-prompt audits are not necessarily wrong, but they may be incomplete: reporting across multiple prompts allows readers to distinguish near-null gaps from substantively meaningful and stable ones.

Statistics

   Article View: 146
   PDF Downloaded: 3