Articles · App ExamplesUpdated September 2026

To add audio recording to an app you need a permission, a file and a plan for both.

To add audio recording to an app you need three things: a microphone permission the user actually grants, a recorder that writes an audio file, and somewhere for that file to live. The recorder is the easy third. The permission wording, the upload queue and the storage bill are the rest of the work, in the same way that adding photo upload turns out to be mostly about the upload.

Below: the exact declaration each platform wants, why only one of them is a sentence you write, what a minute of audio weighs, what to show the user who denies the prompt, and when a voice recording feature is worth building rather than buying.

See what a minute of audio weighs

The short version

The capture code is the small part, the permission and the storage are the project.

Every mobile platform ships an audio recorder and every cross platform framework wraps it. Start, stop, get a file. What none of them hand you is a prompt people accept, a retention rule, an upload that survives a basement with no signal, and an answer for what happens when the phone is wiped.

The limit worth knowing before you start is that a recording is opaque. Nobody skims audio. If the point of the feature is that somebody later finds what was said, you are building transcription with a microphone attached, and that is a bigger product than a record button.

Recording audio is three jobs, not one

Job one is capture. Ask for the microphone, start a recorder, show the user that something is happening, stop it into a file. On a React Native and Expo project that is one library and two buttons, and a working prototype is an afternoon. Nothing in this part will surprise you.

Job two is the file. A recording is a real file on disk with a codec, a duration and a size, and it goes away with the app unless you move it somewhere. That is the same problem shape as any other capture feature, which is why the react native camera work feels familiar the second time you do it.

Job three is the point of the recording. A voice memo in an app that only its author replays needs almost nothing: a local file, a list and a play button. A recording somebody else reviews needs upload, sync, rules about who can hear it and a delete that really deletes. Decide which one you are building first. They share a record button and nothing else.

Two declarations, and only one of them is a sentence

On iOS the Info.plist key is NSMicrophoneUsageDescription. Apple's documentation describes it as a message that tells people why the app is requesting access to the device's microphone, and states that the key is required if your app uses APIs that access the microphone. Your text appears inside the system prompt, so write what you record and what happens to it afterwards, in the words your users would use.

Android has no equivalent sentence. You declare android.permission.RECORD_AUDIO in the manifest and the system writes the prompt wording itself. The asymmetry matters at review time: a vague purpose string is a reason an iOS build comes back, and on the Android side there is nothing to be vague with.

Recording in the background is a third declaration on both platforms rather than a setting. Carrying on while the app is backgrounded or the screen is locked needs the audio background mode on iOS, and on Android a foreground service with a microphone service type plus a notification the user can see the whole time. If you only record while the screen is on and the button is held, you can skip all of that, and most apps should.

Apple, NSMicrophoneUsageDescription

Plan for the user who says no

Android lists RECORD_AUDIO at protection level dangerous. In practice that means declaring it grants nothing. The app asks at runtime, the person answers, and they can take it back later from system settings without telling you. iOS behaves the same way from the user's side.

So microphone access is a state you read, not a fact you learned at install. Check it every time the record button is tapped. The screen needs three states and the third is the one teams skip: never asked, asked and refused, and granted. Skipping the refused state produces a record button that silently does nothing, which users report to you as a crash.

Give the refusal somewhere to go. One line saying what the app cannot do, plus a button that opens system settings, turns a dead feature into two taps. And where audio is optional in your product, let people carry on without it rather than parking them at a wall they already said no to.

Android developers, Manifest.permission reference

What a minute of audio actually weighs

Audio sits between text and video, and people budget for the wrong neighbour. A compressed voice recording at 64 kbps is about 480 KB a minute, so an hour is under 30 MB. The same hour captured uncompressed at 16 bit and 44.1 kHz in mono is over 300 MB. Choosing a voice codec rather than a music one removes most of the storage question before it appears.

The multiplier is people, not minutes. Ten thousand users leaving one two minute note a week comes to close to 10 GB a week at 64 kbps, roughly 500 GB in a year, and it does not stop growing on its own. Pick a retention rule on day one, even if the rule is that nothing is ever deleted, because that is a decision too.

Then assume the signal is gone. Recordings happen in basements, on planes and on sites, which is exactly where uploads fail. The answer is a local queue that retries later, which is the subject of offline app functionality, and without one a recording exists on a single handset and nowhere else. The first person to drop their phone will find that out for you.

Try it

What your recordings weigh

File size is bitrate times duration. Every other storage number on this page falls out of those two.

1.0 MB each, 384.0 MB a month

At 64 kbps. Uncompressed audio is about eleven times the size of a 64 kbps voice recording, which makes the codec the biggest storage decision in the feature.

Five ways to ship it, and what each one gives you

ApproachWorks with no signalNeeds a backendStorage growsGives you searchable text
Local file onlyYesNoon the handsetNo
Local file plus background uploadYesYeson your billNo
Upload, then transcribe laterYesYeson your billYes
Live streaming transcriptionNoYestext onlyYes
Typed notes, no audio at allYesNonegligibleYes

Deciding whether to build it at all

Build it when the recording belongs to something: an inspection note attached to a job, a voice reply inside your own messaging, a field observation that has to land on the right record. Skip it when what people want is a file in a chat app they already use. The honest middle case is the one where a phone's own voice memo app plus a share sheet covers most of the need, and the only real gap is that nobody can tell afterwards which job the file belonged to.

Newly is an AI app builder. You describe the app in plain English and it writes a real React Native and Expo project you own, runs it on a cloud iPhone or Android simulator while it builds, and ships it to TestFlight and to Google Play internal testing. It is $25 a month and there is no free plan. It does not come with a transcription service and it does not choose your backend, so the storage decisions on this page stay yours.

Where this page stops is the moment after the audio exists. A recording cannot be skimmed, cannot be searched, and cannot be used at all by someone who cannot hear it. One thing fixes all three: text sitting next to the audio. That is transcripts and access territory, and it is worth reading before you call the feature finished.

Questions people ask about adding audio recording

Three steps. Declare the microphone permission on each platform, use the platform recorder to capture into a file, then decide where that file lives. On iOS that means the NSMicrophoneUsageDescription key in Info.plist. On Android it means android.permission.RECORD_AUDIO in the manifest plus a runtime request. The capture itself is usually the shortest part of the job.

Describe the recorder your app actually needs

Write down who records, who listens afterwards, and how long the file has to survive. Almost every other decision in the feature falls out of those three answers.

Start building