Much of impact investing still depends on structured, written information. Application forms, due diligence templates and monitoring tools usually assume that people can express themselves clearly in a dominant written language and within predefined fields.
That assumption excludes useful evidence. Entrepreneurs may explain their business more naturally in conversation, while farmers, workers and customers may describe change more accurately in their first language. Programme participants can have detailed experience to share, but little reason to trust a form that compresses that experience into a checkbox.
Automatic Speech Recognition (ASR) offers a route to capture this information at greater scale. A technical research report prepared for Kipimo Solutions by Kevin Obote at Guild Code sets out both the promise and the constraints of building speech systems for Swahili and other East African languages.
Why speech recognition matters for inclusion
ASR converts spoken audio into text. In an impact investment workflow, an enterprise founder could respond to screening questions by voice, a field officer could conduct a semi-structured interview in Swahili, and a beneficiary could describe an outcome without first translating it into formal written English.
Conversational responses can carry context, uncertainty and detail that structured forms often remove. They may also reduce the advantage held by organisations with stronger proposal-writing capacity, better internet access or greater familiarity with investor terminology.
For technologists, the challenge is to make those responses legible without stripping away what makes them useful. A speech system must recognise the words, preserve meaning and record enough context for a human reviewer to judge the evidence.
The current baseline needs local adaptation
Kevin’s report focuses on OpenAI’s Whisper family as a practical open-weight starting point. Whisper large-v3 contains around 1.55 billion parameters and can be fine-tuned on a 16GB consumer GPU using QLoRA, which combines low-rank adapters with 4-bit quantisation.
Zero-shot performance on Swahili remains weak. One African-language benchmark cited in the review places Word Error Rate at about 51% with roughly one hour of adaptation data, improving to about 17% with around 400 hours. The gains begin to flatten after approximately 100 hours.
A system that misunderstands half of the words in a response cannot support a fair screening or impact assessment process. Fine-tuning, careful data selection and honest benchmarking are part of the inclusion design.
Open data and modest compute can take the work forward
The research identifies several open speech resources that can support Swahili and East African ASR, including Mozilla Common Voice Kiswahili, ALFFA, Gamayun Coastal and Congolese Swahili, Kencorpus and AfriVoices-KE.
These datasets provide a credible starting point without requiring immediate large-scale data collection. They also show why one blended label for “Swahili” is insufficient. Kenyan, Tanzanian, coastal, Zanzibari, Congolese and Ugandan variants carry differences that can affect recognition quality.
A dependable system should therefore track dialect, geography and recording conditions where metadata allows. Testing should also examine performance across different speaker groups rather than relying on one aggregate score that hides uneven accuracy.
The report recommends a two-track development path. Teams can test cleaning, augmentation and fine-tuning methods on smaller Whisper models, then validate the strongest configuration on Whisper large-v3 or large-v3-turbo. This reduces iteration costs while preserving a production-relevant final benchmark.
Accuracy needs more than one number
Word Error Rate is the standard ASR metric, but the report recommends publishing both Word Error Rate and Character Error Rate. Character Error Rate can be more informative for morphologically rich languages such as Swahili, where word-boundary mistakes may inflate the headline error without fully reflecting intelligibility.
The report also calls for raw and normalised scores, contamination checks and held-out test data. These choices reduce the risk of overstating performance because test material overlaps with a model’s original training data or because text normalisation quietly improves the reported result.
A transcription score is not proof that the resulting evidence is fair, complete or ready for automated decision-making. Teams must still inspect which speakers, dialects and question types produce the most errors. Kevin’s report is itself a working research review rather than a record of original experiments, so the next stage is empirical testing.
Designing voice into impact workflows
The strongest opportunity is to design ASR around the decision process rather than attach it to an existing form. A useful workflow could allow people to answer in their first language, transcribe the response, preserve the original audio with consent and flag low-confidence sections for human review.
Technologists should also separate recognition from interpretation. The ASR model converts speech into text; a later system may summarise, classify or compare that text against an investment framework. Each stage introduces different risks and should be evaluated independently.
Designed carefully, these layers can widen whose knowledge enters an investment process. Founders can explain business models conversationally, field teams can capture qualitative evidence with less manual transcription, and underserved participants can contribute information in forms that better reflect how they communicate.
The wider promise is an impact investing system that values evidence without requiring everyone to produce it in the same language or format. Better local speech technology can help make that possible, provided accuracy, dialect coverage, consent and human review remain central to the design. Find out more.