SOUVERAIN TECHNOLOGIE
← Back to articles

Voice Dictation for Clinical Records: Running Whisper Locally with Zero Cloud Exposure

By Vincent Savard • • Technology & Healthcare

Voice dictation saves psychologists and attorneys hours of typing every week.

However, consumer applications such as Otter.ai or online cloud APIs transmit live audio to external servers. For psychotherapy sessions or confidential depositions, audio files contain identifiable voices and sensitive disclosures. Routing this data through external APIs violates confidentiality codes and Quebec’s Law 25.

Because OpenAI released Whisper under the permissive MIT open-source license, the entire model can be executed locally without an internet connection.

Hardware Requirements and Benchmark Performance

Running Whisper locally does not require an enterprise data center.

Typical hardware deployed for clinics:

  • Processor: Apple Silicon (M2/M3/M4 with 16 GB unified RAM) or Linux/Windows workstations equipped with a mid-range NVIDIA GPU (e.g., RTX 4060 with 8 GB VRAM).
  • Inference Engine: whisper.cpp or faster-whisper (8-bit quantization via CTranslate2).
  • Model weights: large-v3-turbo or medium.

On an Apple M3 chip, transcribing a 10-minute clinical dictation takes approximately 45 seconds. Accuracy on Quebec French and specialized terminology matches commercial cloud APIs.

Verification by Network Isolation

The core advantage of local deployment is physical auditability:

  1. Firewall rules explicitly drop all outbound network connections for the transcription process.
  2. The software operates identically with Wi-Fi disabled and Ethernet unplugged.
  3. Temporary audio buffers are processed entirely in RAM and wiped once text is transferred to the record.

Want to assess your firm's data sovereignty?

We review your document workflows and provide an actionable independence roadmap with no obligation.

Request a Demo