Articles · App ExamplesUpdated September 2026

To add voice input to an app you pick a recognizer, and the rest is a button.

Every current phone ships with a speech recognizer that any app can call, so to add voice input to an app you ask the operating system for a transcript and put a microphone button next to your text field. The work is a permission prompt, a recording session, a stream of partial results and an editable field at the end. It is the same plumbing as adding audio recording, with one extra decision on top: whether the audio is transcribed on the phone or sent to a server.

That decision is where the honest limits live. On device recognition is free, works with no network and keeps the audio on the handset, but it is less accurate and covers fewer languages. Cloud recognition is more accurate and is billed per minute of audio. Apple's long standing iOS speech API also stops any recognition task that runs longer than one minute, so anything resembling long form dictation has to be chunked whichever route you take.

See what a month of transcription costs

The short version

Use the recognizer the phone already has, and pay only when it is not good enough.

Start with the system recognizer. It needs no account, no key and no server, and it is good enough for a search box, a note field or a spoken command. Ask for microphone and speech recognition permission, start the task, show partial results as the words arrive, and let the user edit the final text before you act on it.

Move to a cloud speech to text service when you need accuracy on long recordings, speaker labels, unusual vocabulary or a language the device does not support offline. That brings a per minute bill, an audio upload and a server of your own to hold the API key. Price it and write it into your privacy copy before you build it, not after.

The phone already has a speech recognizer

You do not need a model, a GPU or an account to start. Both mobile platforms expose a system speech recognizer to third party apps, and the shape is the same on each: request microphone access, request speech recognition access, start a recognition task against either a live audio stream or a recorded file, and read the transcripts that come back. For a search field or a quick note, that is the entire feature.

Apple is unusually direct about the limits, and they are worth reading before you promise anyone a speech to text mobile app. The Speech framework documentation says recognition tasks lasting longer than one minute are stopped, that individual devices may be limited in the number of recognitions performed per day, and that an app can be throttled globally on the number of requests it makes per day. It also tells developers not to send passwords, health or financial data for recognition at all. Those are product constraints, not edge cases, and the last one rules out a voice field on a login screen.

On device recognition is a flag rather than a separate API. On iOS you check whether the recognizer supports on device work, then set the request to require it, and Apple states plainly that on device requests will not be as accurate. The framework now also offers a newer analyzer with a transcriber module that has its own check for whether the current device supports the models it needs. The trade is the one you meet across on device ai mobile apps: nothing leaves the handset, nothing costs money, and quality depends on what the device can hold.

Apple, SFSpeechRecognizer: daily request limits and the one minute audio cap

What cloud transcription actually costs

Cloud recognition is priced per minute of audio and the numbers are small enough to be misread in both directions. Google Cloud Speech-to-Text charges $0.016 a minute for standard recognition on its v2 API up to 500,000 minutes a month, then $0.01, $0.008 and $0.004 a minute in higher bands. Its dynamic batch option, which processes audio at a lower level of urgency, is $0.003 a minute. The older v1 API gives 60 free minutes a month and then charges $0.024 a minute without data logging, or $0.016 a minute with it.

Two lines of the small print change the estimate. Billing is measured in increments of one second and each request is rounded up, so ten thousand three second voice commands are not the rounding error they look like. And each audio channel is billed separately, so a stereo recording of a two person conversation costs twice what the wall clock suggests. Record mono unless you have a reason not to.

Put your own numbers in before you decide anything. Five hundred people dictating ninety seconds a month is 750 minutes, which is twelve dollars at the standard streaming rate. The same five hundred people uploading an hour each is 30,000 minutes, which is $480 a month at the same rate. At small scale, cost is almost never the reason to avoid the cloud. Accuracy, latency and what happens with no signal are the reasons.

Google Cloud, Speech-to-Text pricing

Try it

What a month of transcription costs

Cloud figures are Google Cloud Speech-to-Text list prices for the first 500,000 minutes a month. On device recognition costs nothing and is less accurate.

$12.00 a month

750 minutes of audio at $0.016 a minute. Each audio channel bills separately, so a stereo recording doubles this.

The microphone button is most of the work

Calling the recognizer is the easy part. What makes a dictation feature app feel finished is the second between the tap and the first word on screen. Show a live level meter or a moving waveform so the user can see the app is listening. Print partial results as they stream in, even though they will jump around. Keep the final transcript editable, because it will be wrong sometimes, and a user who cannot fix it quietly stops using the feature.

Voice search in app menus and long lists needs one extra rule: do not run the query on the first partial result. Wait for the recognizer to settle, show the user what it heard, then search. A query that fires on a half heard phrase looks broken even when the transcription was fine, and the user has no way to tell which part failed.

Voice input is also an accessibility feature, and it has to work with the ones already on the phone. A screen reader user needs the microphone button labelled, needs to know that recording has started and stopped without watching an animation, and needs an obvious way to cancel. The same rules that govern every other control apply here, and they are set out in the app accessibility guidelines. One check before you build any of it: on most phones the system keyboard already has a microphone key, so for a plain text field you may be rebuilding something the user has already got.

Decide where the audio goes before you build

Two decisions outlive the button. The first is whether you keep the audio at all. A transcript is small, readable and easy to delete. A recording is a liability with a retention policy attached, and if you store both then you have shipped two features instead of one. Keep the recording only if a person will genuinely need to listen to it later.

The second is where the cloud call is made. An API key cannot live in a mobile app, because anyone who installs the app can read it out of the bundle. So the audio goes to an endpoint of yours first, which turns a screen level feature into a question about where transcription runs and what else that server is responsible for. It is a small server, but it is a server, and it is the point where this stops being an afternoon of work.

Permissions are the last thing to settle and the easiest to leave until review. iOS shows the user a purpose string for the microphone and a second one for speech recognition, and Android asks for microphone access at runtime. Those strings are the only explanation most people will ever read about what you do with their voice. Write one plain sentence, and make sure it is true.

What each route gives a user

RouteWorks with no networkPer minute costHandles long audioWhat you have to build
System recognizer, default settingsdepends on the languagenonechunk ita permission and a button
System recognizer, on device onlyYesnonechunk ita support check and a fallback
Cloud API, live streamingNobilled per minuteYesa server to hold the key
Record now, transcribe laterrecords offlinebilled per minuteYesupload, queue and storage
The keyboard's own microphone keydevice dependentnonethe user decidesnothing at all

Building it into an app you own

If the app already exists, voice input is a library, a permission string and an afternoon. If it does not exist yet, the microphone button is the smallest part of what you are making, and the real question is how to get a real mobile app around it without spending a month on the scaffolding first.

Newly is an AI app builder. You describe the app in plain English and it writes a real React Native and Expo project that you own, runs it on a cloud iPhone or Android simulator while it builds, and ships it to TestFlight and to Google Play internal testing. Plans are $25 a month and there is no free plan. Audio is already in the base project, but a speech recognition package is a native dependency, so adding one triggers a native rebuild that takes a few minutes before the simulator picks it up. After that, every change to the button, the waveform and the transcript screen is ordinary JavaScript and lands in seconds.

The code stays yours throughout. Two way GitHub sync from the Deploy tab creates a repository you own that keeps in step both ways, and Settings exports the project as a ZIP. If the speech library you want needs native configuration that no builder will do for you, take the project out and finish that part in Xcode.

Questions people ask about voice input

Ask for microphone permission, ask for speech recognition permission, start a recognition task on the audio stream, and show the partial transcripts as they arrive. Both mobile platforms provide the recognizer, so for a search box or a note field there is no model to host and no key to manage. The extra decision is whether recognition happens on the device or on a server.

Describe the app the voice button belongs to

Write down what the user says out loud and what should happen next, then build the screen around that answer instead of around the microphone.

Start building