NeFut Logo NeFut
中 Admin Login

[CS.AI] Rethinking Data Quality for AI-Driven Systems: Evidence from Practitioner Interviews

Published at: 2026-09-30 22:00 Last updated: 2026-10-06 12:11
#algorithm #AI #Machine Learning

In AI‑driven software systems, data is no longer just an input; it shapes model behavior, evaluation outcomes, and lawful use. We interviewed 16 practitioners from nine organisations and applied reflexive thematic analysis, extracting six themes. Traceability shifted from modular debugging to attributing model behaviour. Using models to assess data quality introduced a circular dependency. The context and memory of agents became data objects, and synthetic or pseudo‑labeled data raised authenticity concerns. For foundation‑model training, lawfulness acted as a gate, while representativeness was judged by coverage of situations where the system must operate safely. Prior ML work often treats these issues separately, but our practitioner‑grounded account shows they co‑occur as engineering and organisational concerns. We identify five recurring conditions that help explain how these themes erode trust in data and AI outcomes. Finally, we propose a lifecycle‑assurance framing: when data’s influence is embedded in model behaviour, model‑based judgments, or agent actions, evidence must be produced that the data can support a specific AI claim.

Review

Original Source: https://arxiv.org/abs/2609.31191

[h] Back to Home