To add voice input to an app you pick a recognizer, and the rest is a button.
Every current phone ships with a speech recognizer that any app can call, so to add voice input to an app you ask the operating system for a transcript and put a microphone button next to your text field. The work is a permission prompt, a recording session, a stream of partial results and an editable field at the end. It is the same plumbing as adding audio recording, with one extra decision on top: whether the audio is transcribed on the phone or sent to a server.
That decision is where the honest limits live. On device recognition is free, works with no network and keeps the audio on the handset, but it is less accurate and covers fewer languages. Cloud recognition is more accurate and is billed per minute of audio. Apple's long standing iOS speech API also stops any recognition task that runs longer than one minute, so anything resembling long form dictation has to be chunked whichever route you take.
See what a month of transcription costsThe short version
Use the recognizer the phone already has, and pay only when it is not good enough.
Start with the system recognizer. It needs no account, no key and no server, and it is good enough for a search box, a note field or a spoken command. Ask for microphone and speech recognition permission, start the task, show partial results as the words arrive, and let the user edit the final text before you act on it.
Move to a cloud speech to text service when you need accuracy on long recordings, speaker labels, unusual vocabulary or a language the device does not support offline. That brings a per minute bill, an audio upload and a server of your own to hold the API key. Price it and write it into your privacy copy before you build it, not after.
The phone already has a speech recognizer
You do not need a model, a GPU or an account to start. Both mobile platforms expose a system speech recognizer to third party apps, and the shape is the same on each: request microphone access, request speech recognition access, start a recognition task against either a live audio stream or a recorded file, and read the transcripts that come back. For a search field or a quick note, that is the entire feature.
Apple is unusually direct about the limits, and they are worth reading before you promise anyone a speech to text mobile app. The Speech framework documentation says recognition tasks lasting longer than one minute are stopped, that individual devices may be limited in the number of recognitions performed per day, and that an app can be throttled globally on the number of requests it makes per day. It also tells developers not to send passwords, health or financial data for recognition at all. Those are product constraints, not edge cases, and the last one rules out a voice field on a login screen.
On device recognition is a flag rather than a separate API. On iOS you check whether the recognizer supports on device work, then set the request to require it, and Apple states plainly that on device requests will not be as accurate. The framework now also offers a newer analyzer with a transcriber module that has its own check for whether the current device supports the models it needs. The trade is the one you meet across on device ai mobile apps: nothing leaves the handset, nothing costs money, and quality depends on what the device can hold.
Apple, SFSpeechRecognizer: daily request limits and the one minute audio cap
What cloud transcription actually costs
Cloud recognition is priced per minute of audio and the numbers are small enough to be misread in both directions. Google Cloud Speech-to-Text charges $0.016 a minute for standard recognition on its v2 API up to 500,000 minutes a month, then $0.01, $0.008 and $0.004 a minute in higher bands. Its dynamic batch option, which processes audio at a lower level of urgency, is $0.003 a minute. The older v1 API gives 60 free minutes a month and then charges $0.024 a minute without data logging, or $0.016 a minute with it.
Two lines of the small print change the estimate. Billing is measured in increments of one second and each request is rounded up, so ten thousand three second voice commands are not the rounding error they look like. And each audio channel is billed separately, so a stereo recording of a two person conversation costs twice what the wall clock suggests. Record mono unless you have a reason not to.
Put your own numbers in before you decide anything. Five hundred people dictating ninety seconds a month is 750 minutes, which is twelve dollars at the standard streaming rate. The same five hundred people uploading an hour each is 30,000 minutes, which is $480 a month at the same rate. At small scale, cost is almost never the reason to avoid the cloud. Accuracy, latency and what happens with no signal are the reasons.
Google Cloud, Speech-to-Text pricing
Try it
What a month of transcription costs
Cloud figures are Google Cloud Speech-to-Text list prices for the first 500,000 minutes a month. On device recognition costs nothing and is less accurate.
$12.00 a month
750 minutes of audio at $0.016 a minute. Each audio channel bills separately, so a stereo recording doubles this.
The microphone button is most of the work
Calling the recognizer is the easy part. What makes a dictation feature app feel finished is the second between the tap and the first word on screen. Show a live level meter or a moving waveform so the user can see the app is listening. Print partial results as they stream in, even though they will jump around. Keep the final transcript editable, because it will be wrong sometimes, and a user who cannot fix it quietly stops using the feature.
Voice search in app menus and long lists needs one extra rule: do not run the query on the first partial result. Wait for the recognizer to settle, show the user what it heard, then search. A query that fires on a half heard phrase looks broken even when the transcription was fine, and the user has no way to tell which part failed.
Voice input is also an accessibility feature, and it has to work with the ones already on the phone. A screen reader user needs the microphone button labelled, needs to know that recording has started and stopped without watching an animation, and needs an obvious way to cancel. The same rules that govern every other control apply here, and they are set out in the app accessibility guidelines. One check before you build any of it: on most phones the system keyboard already has a microphone key, so for a plain text field you may be rebuilding something the user has already got.
Decide where the audio goes before you build
Two decisions outlive the button. The first is whether you keep the audio at all. A transcript is small, readable and easy to delete. A recording is a liability with a retention policy attached, and if you store both then you have shipped two features instead of one. Keep the recording only if a person will genuinely need to listen to it later.
The second is where the cloud call is made. An API key cannot live in a mobile app, because anyone who installs the app can read it out of the bundle. So the audio goes to an endpoint of yours first, which turns a screen level feature into a question about where transcription runs and what else that server is responsible for. It is a small server, but it is a server, and it is the point where this stops being an afternoon of work.
Permissions are the last thing to settle and the easiest to leave until review. iOS shows the user a purpose string for the microphone and a second one for speech recognition, and Android asks for microphone access at runtime. Those strings are the only explanation most people will ever read about what you do with their voice. Write one plain sentence, and make sure it is true.
What each route gives a user
| Route | Works with no network | Per minute cost | Handles long audio | What you have to build |
|---|---|---|---|---|
| System recognizer, default settings | depends on the language | none | chunk it | a permission and a button |
| System recognizer, on device only | Yes | none | chunk it | a support check and a fallback |
| Cloud API, live streaming | No | billed per minute | Yes | a server to hold the key |
| Record now, transcribe later | records offline | billed per minute | Yes | upload, queue and storage |
| The keyboard's own microphone key | device dependent | none | the user decides | nothing at all |
Building it into an app you own
If the app already exists, voice input is a library, a permission string and an afternoon. If it does not exist yet, the microphone button is the smallest part of what you are making, and the real question is how to get a real mobile app around it without spending a month on the scaffolding first.
Newly is an AI app builder. You describe the app in plain English and it writes a real React Native and Expo project that you own, runs it on a cloud iPhone or Android simulator while it builds, and ships it to TestFlight and to Google Play internal testing. Plans are $25 a month and there is no free plan. Audio is already in the base project, but a speech recognition package is a native dependency, so adding one triggers a native rebuild that takes a few minutes before the simulator picks it up. After that, every change to the button, the waveform and the transcript screen is ordinary JavaScript and lands in seconds.
The code stays yours throughout. Two way GitHub sync from the Deploy tab creates a repository you own that keeps in step both ways, and Settings exports the project as a ZIP. If the speech library you want needs native configuration that no builder will do for you, take the project out and finish that part in Xcode.
Questions people ask about voice input
Ask for microphone permission, ask for speech recognition permission, start a recognition task on the audio stream, and show the partial transcripts as they arrive. Both mobile platforms provide the recognizer, so for a search box or a note field there is no model to host and no key to manage. The extra decision is whether recognition happens on the device or on a server.
Describe the app the voice button belongs to
Write down what the user says out loud and what should happen next, then build the screen around that answer instead of around the microphone.
Start building