Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal
Source: arXiv:2607.11771 · Published 2026-07-13 · By Umm-e- Habiba, Lucas Mauser, Jonas Fritzsch, Justus Bogner, Stefan Wagner
TL;DR
This paper investigates how existing Requirements Engineering (RE) practices support explainability requirements in AI-based systems, focusing on an industrial context. Explainability is recognized as a crucial non-functional requirement in safety-critical and regulated domains, but the authors identify that current RE methods inadequately address explainability needs across the RE lifecycle. Using a multi-phase qualitative study involving eight practitioners from Daimler Truck AG, the authors apply think-aloud protocols and group discussions to explore the elicitation, specification, and validation of explainability requirements. Their findings reveal systematic challenges such as conceptual ambiguity in elicitation, limited expressiveness and testability during specification, and fragmented validation efforts often hampered by vague criteria and regulatory uncertainty.
The study highlights that explainability is not treated as a first-class, continuous concern throughout RE but rather addressed in an ad-hoc and fragmented manner. These cross-cutting, cumulative challenges suggest the need for holistic, process-level support to maintain shared understanding, traceability, and context-aware assessment criteria across RE phases. The authors propose to develop an empirically grounded RE framework that builds on these insights to augment—not replace—current practices by explicitly addressing explainability as a lifecycle-wide concern bridging stakeholder, technical, and regulatory perspectives.
Key findings
- Conceptual ambiguity during elicitation leads to divergent stakeholder interpretations of 'explainability,' resulting in vague or conflicting requirements.
- Explainability requirements were deprioritized relative to more tangible system properties like safety and performance during elicitation activities such as brainstorming.
- Structured Natural Language specifications often overlapped explainability with other non-functional requirements (transparency, fairness) and lacked clear measurability or testability.
- Use Case Diagrams provided limited support for cross-cutting explainability concerns and omitted acceptance criteria.
- Softgoal Interdependency Graphs allowed modeling complex trade-offs but imposed high cognitive load and required early system understanding.
- Validation techniques struggled with vague or underspecified requirements, focusing more on linguistic clarity than on effective user understanding or decision support.
- Compliance checking raised regulatory considerations but was hindered by abstract legal language, confidentiality tensions, and lack of measurable acceptance criteria.
- Challenges encountered in each RE phase compounded one another; early ambiguities propagated through specification and validation, undermining traceability and accountability.
Threat model
n/a — The paper is focused on empirical evaluation of engineering practices within an industrial context rather than adversarial or attacker models. The 'adversary' implicitly refers to challenges emerging from ambiguous stakeholder understanding, regulatory complexity, and method limitations rather than malicious actors.
Methodology — deep read
The study applies a multi-phase qualitative approach focusing on the Requirements Engineering lifecycle steps of elicitation, specification, and validation to investigate explainability requirements for AI-based systems in an industrial setting. The threat model implicitly assumes practitioners as users of RE practices aiming to elicit, specify, and validate explainability amidst challenges like ambiguous terminology, stakeholder diversity, and regulatory uncertainty, rather than a direct adversarial attacker perspective.
Eight practitioners from Daimler Truck AG with 1–10+ years of experience in AI/ML projects and RE processes participated. They were split into groups, each assigned different established RE techniques per lifecycle step drawn from prior literature on AI RE gaps: elicitation (interviews, brainstorming, task analysis), specification (structured natural language, use case diagrams, softgoal interdependency graphs), and validation (structured user feedback, walkthroughs, compliance checking).
The study used think-aloud protocols combined with moderated group discussions across the phases. Participants verbalized their reasoning, challenges, and decisions while working on eliciting, specifying, and validating explainability requirements grounded in a real industrial AI system context (the Active Side Guard Assist feature mandated by EU and UN regulations).
Sessions were audio-recorded, transcribed, and qualitatively analyzed using the constant comparison method guided by grounded theory principles. The researchers iteratively developed and refined coding schemes to identify step-specific and cross-cutting challenges. Requirements from elicitation were consolidated to eliminate redundancy before passing to specification, and similarly into validation. Group discussions enabled participant reflection and triangulation of findings.
No formal quantitative evaluation, metrics, or benchmarks were reported. The study is exploratory and descriptive, focusing on understanding practical difficulties and conceptual gaps in real-world RE practices rather than measuring effectiveness or accuracy. Implementation details like hardware, training epochs, or hyperparameters are not applicable.
The evaluation thus centers on thematic analysis of verbalized participant experiences and synthesis of challenges per RE step. No code or datasets are publicly released. The analysis covers one real-world example scenario end-to-end but is currently preliminary, with later phases planned to develop and evaluate a dedicated RE framework for explainability based on these findings.
Technical innovations
- First holistic empirical study examining elicitation, specification, and validation of explainability requirements across the complete RE lifecycle in an industrial AI system context.
- Use of think-aloud protocols combined with group discussions to capture practitioners’ reasoning and challenges related to explainability requirements engineering.
- Synthesis of step-specific and cross-cutting challenges revealing explainability’s systemic treatment as a fragmented, non-primary concern in existing RE practices.
- Proposal of a research vision to develop an empirically grounded RE framework focused on process-level support for explainability that augments rather than replaces existing techniques.
Datasets
- Daimler Truck Active Side Guard Assist scenario — real-world industrial use case context — proprietary/internal to Daimler Truck AG
Baselines vs proposed
- Interviews (elicitation): conceptual ambiguity reported — challenge identified vs Brainstorming: explainability deprioritized vs Task Analysis: weak goal-function linkage
- Structured Natural Language (specification): overlaps with other NFRs and limited testability vs Use Case Diagrams: limited expression of cross-cutting explainability concerns vs Softgoal Interdependency Graphs: complex dependencies but high cognitive load
- Validation via User Feedback and Walkthroughs: vague, underspecified requirements limit assessment vs Compliance Checking: regulatory considerations raised but legal ambiguity limits applicability
Limitations
- Preliminary phase only; subsequent phases to develop and validate the proposed RE framework are ongoing.
- Small sample size (8 participants), limiting generalizability beyond Daimler Truck organizational context.
- No quantitative evaluation or benchmarking of RE techniques’ effectiveness for explainability requirements.
- Absence of legal experts in validation constrained meaningful regulatory assessment.
- Focus on a single AI system scenario; transferability to other domains or AI applications remains to be studied.
- Dependence on self-reported and verbalized participant data introduces potential biases and subjectivity.
Open questions / follow-ons
- How can a unified RE framework effectively bridge varying stakeholder mental models and terminologies around explainability?
- What measurable criteria or metrics can be developed to evaluate explainability requirements across different RE phases?
- How do explainability RE challenges and solutions generalize across other AI application domains and regulatory contexts?
- What role can legal experts and regulators play in co-developing validation procedures for explainability in RE?
Why it matters for bot defense
Bot-defense and CAPTCHA systems increasingly rely on AI components where explainability can influence user trust, compliance, and risk assessment. This paper highlights that current Requirements Engineering practices fall short in systematically addressing explainability requirements end-to-end, which is crucial for transparent and accountable AI-driven security features. Practitioners in bot defense should note the identified pitfalls such as ambiguous stakeholder expectations and limited testability of explainability needs that could weaken transparency assurances. The call for a process-oriented RE framework underscores the need for integrating explainability considerations early and continuously in security system development cycles to maintain traceability and align technical solutions with regulatory demands. Though the study focuses on automotive AI, the insights on lifecycle continuity and validation challenges are transferable to explainability efforts in AI-powered CAPTCHA and bot detection workflows.
Cite
@article{arxiv2607_11771,
title={ Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal },
author={ Umm-e- Habiba and Lucas Mauser and Jonas Fritzsch and Justus Bogner and Stefan Wagner },
journal={arXiv preprint arXiv:2607.11771},
year={ 2026 },
url={https://arxiv.org/abs/2607.11771}
}