# Local audio implementation

All processing runs in the browser. Files, microphone samples and text are not uploaded to an inference service. Third-party connections fetch public code/model files only. No API keys, subscriptions, accounts or paid inference are configured.

## Run

- `npm install`
- `npm run dev -- --port 3217`: predev builds and copies the worker, FFmpeg, ONNX and Spanish pronunciation assets.
- `npm run build`: prebuild also generates these assets.
- `npm test`, `npm run lint`, `npx tsc --noEmit`.
- After changing `src/lib/audio`, run `npm run build:audio`; the separate audio worker is not rebuilt by Next hot reload.

Keep `public/engines` in deployment output. It is generated, so it is excluded from Git. Do not use static-export hosting without adapting locale selection: the current Next app reads Accept-Language/cookies on the server. This server renders pages only; it never processes audio.

## Architecture and limits

- `definitions.ts`: the 36 tools, supported controls, English/Spanish copy, limits and model disclosures. The UI exposes only implemented options.
- `engine.ts`: FFmpeg worker orchestration, codecs, filters, editing, metadata and cover art.
- `pcm.ts` / `dsp.ts`: PCM, spectral/pitch estimates, sequencing, synthesis and reverb.
- `models.ts`: local Whisper, Kokoro, Demucs and experimental English Chatterbox.
- `Workspace.tsx`: real file input, microphone recording, ordering, progress, cancellation, previews and downloads.
- Generated workers are discarded after each job to release model memory. Transformers models use the browser cache when available. Storage eviction/private mode can require another download.
- 100 MB per input, 150 MB combined; at most 10 minutes combined decoded audio. Microphone capture stops after 2 minutes. These limits reduce memory failures, but do not guarantee operation on every device.
- Supported download codecs: WAV, MP3, FLAC and OGG. Metadata editing preserves audio without re-encoding for MP3/FLAC. Cover editing/extraction targets FLAC.
- Drum instruments and voice-to-instrument timbres are synthesized. No copyrighted commercial sample packs are included.
- Voice changer applies effects to uploaded/recorded clips; it is not live microphone monitoring. Noise reduction uses spectral processing, not a claim of perfect noise removal.
- Vocal tuning is an experimental monophonic overlap-add algorithm, not studio autotune. Speech to Song generates an accompaniment; it is not a generative singing model.
- BPM/key/Camelot are estimates. Hearing Test is a quiet tone generator, not an audiogram or hearing-age diagnosis.
- Visualizer exports in real time through MediaRecorder. It offers two implemented styles; WebM/MP4 availability depends on the browser. Keep the tab visible.

## Language

Preference order: explicit ES/EN cookie, then weighted Accept-Language, then English. No IP geolocation service is needed. Spanish-speaking users abroad keep their preferred language. The selector persists for a year.

The interface language and audio/text language are separate. Whisper tiny is multilingual. Kokoro uses a Spanish pronunciation engine and Spanish voices for Spanish input; English voices use English pronunciation. Signal processing is language independent.

Cloning uses the English Chatterbox model and requires WebGPU. Its multilingual ONNX variant lacks the configuration needed by this browser integration; do not silently process Spanish with the English model. The page explains this restriction in both languages.

## Costs and distribution

Local inference has no per-request provider bill. Hosting, domain, asset bandwidth and user device/network use are separate costs; a universal zero hosting cost is not guaranteed. Large model downloads are disclosed before processing.

The FFmpeg core and pronunciation engines include GPL components. Preserve their licenses and corresponding source/build information when distributing the website. See `public/licenses/audio-engines.html`. The build copies local integration sources and upstream notices alongside the engines. Model weights are downloaded from their upstream public repositories; maintain their licenses if self-hosting.

## Browser verification

`scripts/browser-audio-smoke.ts` runs real workers against generated PCM fixtures and exercises bilingual TTS/transcription, cloning and video recording/extraction.

Build the test module with `npx esbuild scripts/browser-audio-smoke.ts --bundle --format=esm --outfile=public/engines/smoke.js`, copy `scripts/smoke-page.html` to `public/engines/smoke.html`, open that URL and run the buttons. Model tests download large assets and can take minutes. Remove the two generated smoke files before deployment.

Output generation checks do not establish studio-level quality or compatibility across all browsers. Validate on target mobile devices before a public release.
