BETA

The service is in beta. You may see slight delays — thank you for your patience. Hit a problem? Write to us

Engineering

Why an hour of audio comes back in minutes

Updated 20 September 2026
6 min read
Blogבעברית
Transcription speed

An hour of audio should not cost you an hour of waiting. The recording is split into segments, decoded at the same time, separated by speaker alongside that, and merged with timestamps at the end. The result is that an hour-long meeting normally comes back in minutes — not in seconds, and not in hours. Here is what happens in between, and which parts you can actually influence.

How it is put together

One long recordingSegment 1Segment 2Segment 3Segment 4Segment 5Segment 6Merged with word-level timestamps, diarization running alongsideEvery segment is decoded at the same time. A stuck segment is retried or split in two; none is ever dropped.
Why two hours do not take twice as long as one.
  • Capacity follows demand. Processing workers start when there is a queue and stop when there is not, in three tiers — free, Pro and business. When many recordings arrive at once, more workers start.
  • Segments in parallel. A long recording is cut into segments decoded simultaneously and merged with word-level timestamps. Silence and background noise are filtered before decoding.
  • No segment is dropped. A segment that stalls is retried, and split in two if it has to be. A meeting comes back whole or not at all — never with a hole in the middle.
  • Diarization in parallel. Speaker separation runs beside transcription on the same worker, so it does not stack onto the wait.

What actually decides your wait

FactorHow much it mattersWhat you can do
The queueMost of the variance. A cold start costs a minute or two; a warm worker starts immediately.Nothing — and it is usually invisible.
Length of the recordingMuch less than you would expect, because of the parallelism.Trim the ten minutes of small talk before the meeting started.
File sizeReal, but it is upload time, not processing time.Upload audio rather than video; the limit is 500MB either way.
The summaryA short additional step after the transcript.Re-summarising later reuses the transcript and costs no new minutes.

What it feels like in practice

  • A one-hour meeting: transcript, speakers and summary normally back within minutes.
  • A two-hour lecture: the same order of magnitude, not double.
  • A batch of recordings at once: capacity grows, within your tier's allowance.

The transcript appears before the summary, so you can start reading while the rest finishes — and you can close the tab entirely; processing does not depend on your browser staying open.

Speed is not the thing we optimise

A transcription service can be made arbitrarily fast by accepting worse text. We choose the model that makes fewer errors on real meetings, not the one that finishes first; speed comes from how the work is arranged, not from cutting corners on the decode. How the pieces fit together is described in how IvreetMeet is built.

Try it with a long one

The free plan covers 7 meetings and 120 minutes per organization per calendar month. A single long recording is a better test of this than five short ones.

Frequently asked questions

Does a two-hour recording take twice as long as one hour?

No. The segments are decoded at the same time, so doubling the length does not double the wait.

Why is it sometimes slower?

Usually the queue. When capacity has to grow, there is a minute or two before processing starts; a worker already running begins immediately.

Does speaker separation add time?

Not meaningfully — it runs alongside transcription rather than after it.

Can I make it faster?

Upload audio rather than video when you have the choice: the transcript is identical and the upload is a fraction of the size. On a large file, most of the wait you feel is your own connection.

Try it on your own recording

IvreetMeet transcribes and summarizes Hebrew meetings. The free plan covers 7 meetings and 120 minutes per organization each month.

Start free