Audio to text

Upload a recording and get the text. An interview, a lecture, a call or a voice memo: transcription runs through a model that holds up against noise and accents.

Drag here or paste from the clipboard (Ctrl+V)

0 of 1 selected

mp3, wav, m4a, ogg — up to 20 MB. Drag it here or paste from the clipboard

If you transcribe regularly

Konspekt: meeting notes that never leave your computer

A smart notepad for meetings. Konspekt records the conversation, transcribes it and turns it into notes entirely on your own machine. Nothing is uploaded anywhere: not the audio, not the transcript, not your notes.

Download Konspekt

Free. Installs in a minute and works offline afterwards.

The summary: decisions, owners and open questions
The transcript: every line attributed to a speaker
Your own notes next to the transcript
Dictation: the text lands wherever your cursor is

The summary: decisions, owners and open questions

Notices the conversation

If people are talking on your computer and recording is off, Konspekt offers to start. Nothing is saved until you agree.

Records both sides

Your microphone and the system audio, so the other person in a call is captured too.

Tells speakers apart

Name someone once and Konspekt recognises their voice in later meetings.

Writes the summary

Decisions, who does what, and what is still open — from the actual conversation.

Runs on your machine

Self-hosted in the literal sense: the audio, the transcript and the notes sit in a folder on your disk. Nothing goes to someone else’s server, and the app keeps working offline.

Free, no account

No sign-up, no API key, no subscription. Download it, install it in a minute, use it.

Open source

The code is on GitHub: read it, change it, build it yourself. FSL-1.1-MIT only forbids repackaging it as a competing commercial product.

Learn more

Open source under FSL-1.1-MIT. Every release becomes plain MIT after two years.

What transcription is

Transcription turns spoken words into written text. What was said in a conversation, a lecture or an interview becomes a document you can search, quote and edit.

Upload a file and the model recognises the speech. Depending on the length and the quality of the sound it takes from a few seconds to a few minutes. Accuracy depends on the original recording: noise and overlapping voices hinder a model the same way they hinder a person.

Who needs it

  • Journalism: interviews and press conferences.
  • Education: lectures and seminars as text.
  • Law: hearings and witness statements.
  • Medicine: dictation and consultation records.
  • Business: meetings, webinars and presentations.
  • Video: subtitles and a text version of a clip.