On-Device AI Transcription

& Translation for macOS

Turn your audio and video files into polished transcripts

and subtitles—100% offline, private, and powered by Apple Silicon MLX.

Transcribe and translate any audio or video file with maximum accuracy using 100% local AI models on your Mac. Powered by Apple Silicon MLX, WaveCap runs entirely offline, delivering studio-grade accuracy, instant speaker diarization, and effortless exports to SRT, VTT, TXT, and JSON without cloud fees or privacy risks.

Why You'll Love It

Local and private

100% Local & Private

Processing happens entirely on your Mac. Your audio files and transcripts never leave your device or touch a cloud server.

Blazing speed

Blazing Speed via Apple Silicon

Built natively using Apple's MLX framework to leverage the full power of your Mac's Neural Engine and Unified Memory.

AI polishing and translation

Smart AI Polishing & Translation

Go beyond raw text. Built-in LLMs remove filler words ("um", "uh"), fix grammar, and translate transcripts locally.

Noise cleanup and separation

Noise Cleanup & Separation

Isolate clear vocals from noisy recordings before transcribing using local vocal separation and Voice Activity Detection (VAD).

Interactive Feature Breakdown

Drag and drop batch queue

Drag-and-Drop Batch Queue

Drop single files or entire folders (MP4, MOV, MP3, M4A, WAV, AAC, and more). WaveCap queues and processes them automatically in the background.

Speaker diarization

Speaker Diarization

Automatically detect, separate, and label different speakers throughout your recordings.

Pro-grade exporting

Pro-Grade Exporting

Export clean text (TXT), subtitles (SRT, VTT with optional speaker tags), or structured data (CSV, JSON) formatted for player compatibility.

Powered by Leading On-Device AI Models

WaveCap offers flexible, configurable pipelines powered by open-weights models optimized for Apple Silicon:

Pipeline Step

Supported Models & Runtimes

Speech-to-Text (STT)

Whisper (Tiny to Large v3 Turbo), Parakeet Series, Qwen3 ASR, Cohere Transcribe, Voxtral Mini 4B, SenseVoice

Text Polishing & Translation

Gemma 4 (E2B / E4B), Qwen 3.5 (4B / 9B), Nemotron 3 Nano 4B

Vocal Isolation & Cleanup

HDemucs v3, HTDemucs (Standard / Fine-tuned / 6-source), MDX Series, DeepFilterNet (v2 / v3)

VAD & Diarization

Silero VAD (v5 / v6), Sortformer Streaming 4SPK v2.1

The Engine Behind WaveCap: Built for Apple Silicon with MLX

What is MLX?

MLX is Apple Silicon's official open-source machine learning framework, engineered specifically for M-series chips (M1 through M4 and beyond). Unlike legacy AI frameworks adapted from cloud servers, MLX taps directly into Apple's unified memory architecture and Neural Engine. It acts as a lightweight, lightning-fast engine capable of running state-of-the-art open-weights models (like Whisper, Qwen, and Gemma) right on your Mac.

Why WaveCap Uses MLX

Unified memory advantage

Unified Memory Advantage

Cloud-focused frameworks force models to constantly copy data between the CPU and GPU. MLX takes advantage of Apple's shared RAM architecture, allowing heavy models (like 9B LLMs or 1.7B Whisper variants) to run with minimal overhead and zero memory duplication.

Blazing fast on-device inference

Blazing Fast On-Device Inference

MLX delivers peak single-stream performance. By running model operations directly on the Mac's GPU and Neural Engine, WaveCap transcribes and polishes multi-hour files in a fraction of the time without overheating your machine.

100% offline and zero subscription

100% Offline & Zero Subscription

Because MLX handles model runtimes locally, WaveCap doesn't need to send your private audio to remote API endpoints like OpenAI or Anthropic. You get server-grade accuracy without recurring monthly fees or internet dependency.

Low power and battery efficient

Low Power & Battery Efficient

Engineered to follow Apple's native power management guidelines, MLX allows WaveCap to run heavy batch queues in the background without draining your MacBook's battery or spinning up loud cooling fans.

Frequently Asked Questions

General & Privacy

Do my audio files or transcripts ever get uploaded to the cloud?

Does WaveCap require an active internet connection to work?

How does WaveCap protect my data compared to web-based transcription tools?

Hardware & System Requirements

What are the system requirements for running WaveCap?

How much storage space do the AI models require?

Will running heavy models slow down my Mac or drain the battery?

Features & Workflow

What media formats are supported?

Can WaveCap distinguish between different speakers?

What subtitle and text formats can I export?

How does the AI text polishing work?