Can AI listen better than a human?
AI is taking over more and more tasks in society, and with each one the same question returns: what changes, and what stays distinctly human? A common answer is that relationships remain ours. Listening to each other, showing empathy and compassion, building real connection, these are the things a machine supposedly cannot take from us.
The leonardo team works in the social impact sector, a field built on improving the livelihoods of people around the world, especially those in marginalised groups. Listening to those people is the core of that work, which makes the question above a practical one for us rather than a philosophical exercise.
On 8 September 2026 we released voi, the first AI impact researcher, together with our partners Impact Hub Network, Impact Hub Trento and the Technical University of Munich. voi runs AI-led interviews at scale and, as our launch message puts it, helps you listen better.
That claim deserves scrutiny, and answering it properly means starting with how impact data gets collected today, and what the research says happens when AI leads the process.
Depth or scale: The choice every research team makes
Whether it is a master's thesis, a PhD or a study inside a company, the choice between qualitative and quantitative methods comes down to a trade between depth and scale. Qualitative research goes deep. A researcher interviews a small number of people, follows their reasoning and picks up on what they did not expect to hear. Running those interviews and then coding what people said takes time, methodological rigour and a good deal of emotional intelligence. Quantitative research goes wide. A survey with fixed questions and answer options can reach hundreds of participants and is easy to code, but respondents can only say what the answer options allow them to say.
What organisations do at the moment, and what AI changes
Hundreds of trained researchers who speak several languages, months of time and a large budget are out of reach for almost every social impact project. Faced with that gap, impact organisations tend to do one of four things:
- Skip impact measurement and management as a whole
- Focus solely on output data
- Use a simple and scalable survey to collect quantitative outcome data
- Collaborate with a university or a research institute if the budget is there
This is why we built voi: The first AI impact researcher is capable of running hundreds of interviews at the same time, respondents answer open questions in their own words and by voice, and voi asks intelligent follow-up questions based on what it has just heard. Qualitative impact depth becomes something an organisation can reach at the scale of a survey.
Listening as a human: Where human-led data collection breaks down
Ask an impact team what makes their beneficiary data trustworthy and the answer usually points to the field team, to trained enumerators working in the local language over weeks on the ground.
The person asking changes the answer
Participants adjust their answers towards whoever is in front of them, most of all when that person represents the organisation controlling access to something they want. Two findings set the scale:
- Self-reported illicit drug use runs roughly 1.3 times higher when a questionnaire is self-administered instead of read aloud by an interviewer (Tourangeau & Yan, 2007).
- People who believed their interviewer was automated disclosed more than those who believed a human was behind it, although both groups spoke to the same system (Lucas et al., 2014).
The second effect is quieter and more expensive. Enumerators push answers in different directions from one another, and that variance behaves like a cluster design effect. Brunton-Smith et al. (2017) found interviewer correlations small enough to look negligible producing design effects of two to three at the workloads typical of field surveys, since the penalty scales with how many interviews each enumerator conducts. At three, a 380-household survey carries the statistical information of roughly 125. The instrument the sector trusts most carries an error term it almost never measures.
The economics of listening
Interviews are expensive in a way that shapes the research before it starts. A researcher can only speak to one person at a time, and every hour of conversation is followed by several hours of transcribing and coding the answers. The number of people an organisation can interview is therefore decided by how many staff hours it can afford, and rarely by how many voices the question actually needs.
Most organisations respond by splitting their research in two. They run a survey when they need to reach many people and a small set of interviews when they need depth, and end up with two datasets that were collected separately and cannot be combined into one picture.
A conventional beneficiary-voice study from an established provider runs into the low five figures, out of reach for most social enterprises, mid-sized NGOs and foundation grantees. The qualitative component gets cut, and a closed instrument remains.
The language barrier
A written questionnaire asks people to read and write in the language it was drafted in, and an interviewer-led one asks that someone can reach them and probe fluently. The respondents who clear those hurdles are more educated, more urban and more often men, and nothing in the report records who was missing.
What changes with an AI interviewer
The evidence base is young, and the most detailed study to date (Chopra & Haaland, 2023) interviewed online panel respondents in a high-income country about their financial decisions, so transfer to programme settings is a hypothesis to test rather than a result to assume. Three findings are consistent enough to plan around:
- Consistency: Annotators hand-coded over 18,000 AI-generated interview questions and rated more than 95% as open-ended, neutral and relevant. With one interviewer instead of twelve, there is no between-interviewer variance left to correct for.
- Probing: A skilled human interviewer probes at least as well, and direct comparisons report similar quality ratings, so the difference is one of cost. Measured against a single open-ended survey question, which is the realistic alternative for anyone who could never fund a researcher, AI-led interviews surfaced roughly five times as many distinct themes, most of the gain traceable to follow-ups generated in the moment.
- Scale: Interviews run in parallel rather than one at a time, and each additional one costs almost nothing once the guide exists. They still reached thematic saturation at around 25 respondents, the same range as human-led studies, so the reach comes without thinning what each person says.
What it still cannot do
- Reach. A link cannot get to households without a device or a connection, and that excludes rural households and women first (GSMA, 2026).*
- Observation. A human researcher in the room reads body language, tone and hesitation, and notices what the respondent never thought to mention.
*combining voi with a human led telephone interview is a good way to combine offline mode with the pre- and post- work of the interview.
How voi works: Qualitative depth at the scale and cost of a survey
voi (voices of impact) is our answer for holistic impact data collection and analysis. It works in three steps.
- You describe your programme and the change you want to understand. voi builds a ToC and drafts a scientific interview guide with validated indicators in minutes, which you can edit or align to a framework you already use.
- Your stakeholders speak for themselves, by voice, video or text, in their own language, with questions read aloud so that literacy is not a condition of taking part. voi asks its own follow-up questions across hundreds of conversations in parallel.
- Responses are coded against the indicators the guide was built from, producing values you can report, disaggregate and compare across sites or over time, alongside key themes and an executive summary.
Step three is where the method has to earn trust, so every insight links back to the recording it came from, and you can check what the system reports against what the person said.
The limits above shaped the design as well. Language coverage is treated as a variable to be tested, theme discovery stays supervised, data sits on GDPR-compliant servers in Europe and the system is built to EU AI Act standards. The underlying research is developed with TU Munich, and a study starts at €299.
Listen to the communities you serve, at breadth and depth
Everyone your project, programme or initiative affects can speak for themselves, in their own words and their own language, and what they say becomes evidence you can report, disaggregate and act on, in one study that carries the reach of a survey and the depth of an interview.
And now to the moment you have been waiting for: Can AI listen better than humans? Our answer is no! But it is a great tool to empower organisations to combine depth and scale to gather data that helps to prove and improve the impactful work done all around the globe.
Book a demo and we will run voi on your own programme.
Want to know more?
Get in touch with us and and start to measure impact confidently.