Why your students' recitations should never touch cloud AI
There is a version of this product that would be easier to sell. A teacher records a child reciting, the audio goes to a speech model, and a score comes back: accuracy, tajwīd, fluency, a number out of a hundred. It demos beautifully. Every parent evening would end with somebody asking why we do not have it.
We do not have it, and it is not a roadmap item waiting for engineering time. It is a decision, and this is the reasoning behind it.
What sending audio to a model actually involves
"Cloud AI" is a soft phrase for a specific sequence of events. A recording of a named child, made in a mosque, is uploaded to a company the mosque has no relationship with, stored on infrastructure in a jurisdiction the mosque did not choose, processed by a system the mosque cannot inspect, and retained under terms the mosque did not negotiate and usually has not read.
Ask the concrete questions and the picture sharpens:
- How long is the audio kept after the score comes back?
- Is it used to improve the model? Is that opt-out, opt-in, or neither?
- Who inside that company can listen to it? Under what process?
- What happens to it if the company is acquired, or fails?
- Which country is it in right now, and which law applies to it there?
A mosque with fifty children cannot answer any of those, and cannot realistically negotiate the answers. So the decision is made by default — which is to say it is made by whoever wrote the terms of service.
Why a child's recitation is not ordinary data
Voice is biometric. That is not an activist framing; it is how GDPR Article 9 and most comparable regimes treat it, and it is why voice sits in the same category as fingerprints rather than in the same category as an email address.
Layer on the rest. The speaker is a child. The recording is of religious practice. The consent is given by a parent who is being asked to trust a mosque, not a software vendor, and who will not be shown the sub-processor list. And the recording is, unavoidably, a record of a child being bad at something in front of their teacher — which is precisely the material a family has the strongest interest in keeping inside the room where it happened.
Note: The question is not whether a particular provider is trustworthy today. It is whether a mosque should be in the position of having to assess that, on behalf of families who were never asked.
The feature is not as good as it demos
Set the ethics aside for a paragraph and the engineering case is weaker than it looks.
Automatic scoring of Quranic recitation is genuinely hard. Tajwīd is a set of rules about articulation, elongation and assimilation that general speech models are not trained on and do not represent. Children's speech is recognised substantially less accurately than adults' by nearly every published system. Add a mosque's acoustics, regional pronunciation, and a seven-year-old's volume, and confident-looking output starts to arrive that is simply wrong.
Wrong output from a machine is more damaging than no output, because it carries authority. A teacher who hears a child recite knows what they heard. A number that says 78% invites a parent to argue with a teacher about a child, using a figure neither of them can interrogate.
What Mudrus does instead
Review is written by the teacher. A recitation is recorded, attached to an exact ayah range, and marked word by word by the person who heard it. Accuracy is tracked over weeks, so improvement is visible — but the judgement is a human one, made by someone accountable to that family, in that mosque.
The specifics, so this is checkable rather than reassuring:
- No speech-recognition provider receives any recording. Not a paid one, not a free one, not one of ours. Audio is not in that path at all. - Nothing trains a model. No student data trains anything, ours or anyone else's. - Recordings are optional and consent-gated. They are off until a guardian consents, and withdrawing consent deletes them. A mosque that would rather not record at all loses no other part of the product. - Storage is encrypted, in the region the mosque picks at signup. On desktop and mobile the local database is SQLCipher-encrypted with a key held in the operating system's keychain; a device that cannot encrypt is refused a local store rather than quietly given a plain one.
The part we are not claiming
On-device speech recognition — a model that runs on the teacher's own laptop or phone, with nothing leaving it — is designed and not built. We have written the architecture for it. It is not in the product today, and until it is, the word "transcription" in Mudrus means a teacher typing.
We would rather say that plainly than let the word do work it has not earned. A product that is careful about a family's data and careless about its own claims is only careful in one direction.
The trade we are making
We lose a demo. Somebody evaluating Mudrus against a competitor will see a scoring feature on one side of the comparison and not the other, and some of them will choose the other.
What a mosque gets in exchange is an answer to the question a parent eventually asks: where does my child's voice go? The answer is short, and it does not require anyone to trust a third party they have never heard of.
It stays here.