AI Audio Post

Powered by AIRun by humans

Your video looks finished. It does not sound finished. We take the audio the rest of the way.

One frame, five layers

Two shots, and neither of them arrived with any sound on it. These are the layers that go underneath, and you get every one of them as its own file.

Press play. All five layers, at the level they sit at in the film.Now click any track to hear it on its own. Click it again for the whole mix.

Each one is drawn from the finished mix, at the level it actually sits at. The crowd is on the pavement with the camera, so it is loud. The traffic is a block away, so it is the quietest thing in the mix.

Sound is the tell

Generated picture has caught up fast. Generated sound has not. Dialogue sits flat and slightly wrong. Room tone changes at every cut. Nothing in the frame makes the noise it should make, and the silence where a footstep belongs reads as fake long before anyone can say why.

An audience registers all of it in the first ten seconds. Most of them will never be able to name it. They just stop believing what they are watching.

Why a repair tool cannot do this

Audio repair software opens a sound file. It cleans what is on it: noise, hum, clicks, a voice recorded too quietly. It is good at that, and that kind of processing is part of our engine too.

It has never seen your video. It cannot know a door shuts on frame 412, because the door is not in the file it opened. All it can do is improve a sound you already have.

Most AI video has no sound to improve. The footsteps, the door, the room, the traffic outside were never recorded, and every one of them has to land on an exact frame.

AAPE opens the video. It finds the cuts, the gaps, and the frames where something happens, and the sound is built against the picture from there. Cleaning is the small part of the job.

What you send, what you get

Any video. A short, a feature, an ad, a trailer, a game cinematic, an explainer, a music video, a clip for social. If it has picture and it needs a soundtrack, it is the same work and the same engine.

01 / You send

Your cut

A link to the picture, plus whatever audio already exists. Generated voices, a scratch mix, a temp score, or nothing at all. Any of it works.

02 / We handle

The whole soundtrack

Dialogue repair, noise and room tone, Foley, sound design, ambience, music placement, and the final mix.

03 / You get

Delivery files

A finished mix, separated stems, and files in the formats a platform, a client or a festival will accept without sending them back.

The pipeline

Every stage runs AAPE and an engineer at once. What changes between them is the share, never the fact, so each one is listed with what the engine does and what the engineer decides.

  1. 01

    Intake

    The cut is conformed and logged. Dialogue is transcribed and aligned to picture, scene boundaries are marked, and every spot where sound should exist and does not is written down. This list is what the rest of the work is measured against.

    automated

    Transcription, alignment to picture, scene detection, level and loudness read off the file.

    human

    Watches it end to end and decides what the piece needs, which is never the same list twice.

  2. 02

    Restoration

    Broadband noise, hum, clicks, clipping, sibilance, the digital artifacts that generated voices leave behind, and level inconsistency between takes and between scenes.

    automated

    Our restoration engine, across hundreds of clips without drifting. Measurably steadier than hand work, which is why it does this part.

    human

    Chooses the treatment and how hard it runs, then listens back. Restoration pushed too far takes the voice with the noise, and no meter shows that.

  3. 03

    Build

    Foley, hard effects, backgrounds and ambience against picture. Footsteps, cloth, props, doors, room tone, weather, traffic, the layer that generated video never produces.

    automated

    Searching and auditioning libraries, conforming takes, holding layers in sync.

    human

    Every choice and every position. Timing against a specific frame is the thing no tool does convincingly yet, and it is what the audience hears.

  4. 04

    Mix

    The balance. Dialogue is made intelligible first, then music and effects are placed underneath it. The mix is checked on small speakers and headphones before it is signed off.

    automated

    Dynamics, level matching, and loudness measured against the target while the mix is built.

    human

    The balance itself, by ear on calibrated monitors. What should be loud is a decision, not a measurement.

  5. 05

    Deliver

    Mix, stems and platform masters, measured against the published specification for the destination and corrected if they miss. Nothing goes out without that check.

    automated

    The measurement, file by file, against the published numbers.

    human

    Listens to the master before it leaves. A file can pass every number and still be wrong.

What we need from you

A link to the picture lock, or the closest thing to it you have. Re-cutting after the mix means re-doing the mix, so the later the picture changes, the more it costs.

Any audio that already exists, in whatever state. Generated dialogue, a temp score, production sound from a phone, all of it is useful and none of it is required.

Where it is going. A client, a streaming platform, a festival, a broadcaster, an app store, or the internet. The destination decides the delivery specification, and the specification decides how it gets mixed.

Turnaround

Quoted per project when you send the cut, because runtime is only one of the things that decides it. A ten minute piece with clean dialogue already in it and a five minute one with no audio at all are not the same job.

You get a date with the price, before anything starts.

What we handle

Dialogue

Repair and edit

Cleaning up generated voices, removing the artifacts that give them away, evening out level and tone between lines, and cutting the performance so it sits naturally against picture.

Restoration

Noise and room tone

Broadband noise, hum, clicks and clipping. Matching room tone across shots so a scene assembled from separate generations sounds like one room instead of six.

Foley

Footsteps, cloth, props

Recorded and placed against picture. The layer an audience never notices when it is there and never stops noticing when it is missing.

Sound design

Hard effects and the built ones

Everything the frame implies and nothing produced. Impacts, mechanisms, creatures, weapons, weather, and whatever the picture needs that does not exist yet.

Ambience

Backgrounds

The bed under every scene. Rooms, streets, weather, crowds. It is what makes a cut feel like a place rather than a shot.

Music

Placement and edit

Fitting score or licensed tracks to picture, editing them to land on the cut, and balancing them so they support dialogue instead of fighting it. Original score is available and quoted separately.

Mix

The balance

Every element in one place, mixed by an engineer for the destination. Stereo, 5.1 or immersive depending on where it is going.

Delivery

Files that get accepted

Masters, stems and platform specific files, measured against the published specification for the destination before they go out.

People finish it

AAPE does a real share of the work on every project here and does it well. What it does not do is decide: placement, timing, what should be loud, and whether a technically clean take still sounds wrong.

That part is done by engineers with twenty years of experience. It is not a fallback for when the software fails. It is the reason this works.

What each one does

These two columns are not two halves of the pipeline. They run in the same stage, on the same project, at the same time. Anyone telling you the first column is empty is selling craft nostalgia, and anyone telling you the second one is empty is selling a plugin.

What AAPE does

  • Broadband noise, hum and hiss removal, consistently and without artifacts.
  • Clicks, crackle, clipping and digital faults in generated audio.
  • Transcription and alignment of dialogue to picture.
  • Separating a mixed track back into dialogue, music and effects.
  • Level matching across hundreds of clips without drifting.
  • Measuring a finished mix against a loudness specification.

What a person decides

  • Deciding what the audience should be listening to in any given second.
  • Placing a footstep so it lands on the frame the foot lands on.
  • Knowing that a technically clean take still sounds wrong, and why.
  • Building a sound for something that does not exist and has no reference.
  • Pacing. When the soundtrack should drop out entirely.
  • Taste. Twenty years of hearing what did not work.

Why this is not a platform

There is no upload box here that returns a finished mix in four minutes. Those exist and they are useful for a rough pass, and if that is all your video needs then you do not need us.

What arrives here is usually something somebody cares about, at the point where the one-click pass has already been tried and something is still wrong with it. Fixing that is a conversation and a set of decisions, not a queue job.

Send the cut. You get a price and a date back.

Get a quote