AI and data quality: a growing risk
One of the biggest challenges AI has introduced to market research is a growing data quality challenge. And it’s something the industry is still working through.
Synthetic respondents: the hidden threat
AI tools can now generate convincing open-ended responses, mimic natural language patterns, and produce survey answers that look, on the surface, entirely plausible.
These synthetic respondents, such as bots and AI-generated personas, are increasingly capable of passing standard quality checks, entering research studies, and adding responses that may appear convincing, but aren’t based on real people or real experiences.
This isn’t a theoretical concern. It’s already happening. And its implications for qualitative research, where every participant’s voice matters, are serious. Explore this further in our blog: Synthetic Respondents in Market Research.
The scale of the problem
According to Wave 1 of the Global Data Quality Initiative’s benchmarking study, a meaningful proportion of survey data is routinely removed during fieldwork due to fraud, duplication, and poor-quality responses — typically in the region of 10–15%, with additional post-survey cleaning increasing that figure further.
Why Data Quality Risks Are Growing
Earlier GDQ findings suggest the true scale of problematic data may be significantly higher when inattentive and disengaged responses are taken into account. Together, these findings highlight a critical point: even with active quality controls in place, a substantial proportion of research data cannot be taken at face value.
UK benchmarking has also shown that risk is not evenly distributed. Agency-led projects and supplier-sourced data can carry different levels of exposure, with some supplier routes seeing significantly higher rates of fraud or quality issues. The message is clear: data quality cannot be assumed.
As Debrah Harding, Managing Director of MRS, notes, while the industry should embrace the opportunities AI brings, maintaining ethics, quality and integrity must remain non-negotiable if research is to retain its impact.
It’s not just fraud
Alongside deliberate fraud, AI adoption has amplified other data quality risks. Participant fatigue, inattentive responding, and poorly designed AI-generated screeners that fail to properly filter audiences all contribute to datasets that may appear complete but lack the depth and authenticity research depends on.
Why qualitative research is especially vulnerable
In quantitative research, poor responses can sometimes be statistically absorbed. In qualitative work, where projects may involve 8, 12, or 20 participants, a single fraudulent or disengaged respondent can noticeably distort findings.
Increasingly, this also includes participants using AI tools to help shape or generate responses – not always fraudulently, but in ways that can still impact authenticity and depth. The stakes are higher, and the human safeguards matter more.
How Angelfish approaches data quality
At Angelfish Fieldwork, we have signed the Global Data Quality Excellence Pledge. Every participant we recruit is validated in-house by a trained, RAS-accredited team member.
We speak to people directly, confirm eligibility, and assess articulation and engagement before any fieldwork begins. Read more about our approach on our market research data quality page.