BETA

The service is in beta. You may see slight delays — thank you for your patience. Hit a problem? Write to us

Engineering

How IvreetMeet is built

Updated 20 September 2026
8 min read
Blogבעברית
The IvreetMeet processing pipeline

IvreetMeet does not forward your recording to somebody else's transcription service. Behind every upload is our own speech model, trained on spoken Hebrew and running on our own infrastructure, with speaker separation beside it and a structured summary after it. This is what that arrangement buys you: accuracy, privacy, and time — in that order.

Our own model, for the Hebrew people actually speak

General models trained mostly on English handle clean read-aloud Hebrew and then fail exactly where a meeting lives: interruptions, two people at once, English terms dropped into the middle of a Hebrew sentence, names, numbers. Ours was trained on spoken Hebrew, and it keeps training. Every new version is compared blind against the version currently serving customers, on the same recordings, and only replaces it if it makes fewer errors.

On the public eval-d1 benchmark published by ivrit.ai our model is level with the best public system, by our own measurement. On our own meeting sets, in the most recent comparison, it made fewer errors than its predecessor on every set — and the largest gap was in the hardest audio: room echo, background noise, compression. In other words, in real meetings. Most of that improvement was fewer substituted words, which is the error that changes meaning. That comparison is re-run every time the model changes, and a version that is not better does not ship.

What happens to a recording

What happens to a recording, left to rightUploadaudio or video,or record in the browserTranscriptiona model trained on spokenHebrew, running on a GPUDiarizationwho said what;names are keptSummaryfrom text only: key points,decisions, action itemsExport and sharePDF, Word, PowerPointand share linksWhat is left afterwards, and whereThe processing worker: nothing · Your account: the recording and the transcript, until you delete them · Model training: noThe language model receives text only, never the recording · Recognising a speaker across meetings: by consent only
How a meeting moves through the system. This is its shape, not a measurement.

Four things worth pulling out of that diagram:

  • Transcription is parallel. The recording is split into segments decoded at the same time and merged with word-level timestamps, which is why two hours do not take twice as long as one. Detail in why an hour comes back in minutes.
  • Diarization runs alongside, not after. Speaker separation happens on the same worker while transcription proceeds, so it adds no meaningful wait.
  • Capacity follows the queue. Workers come up when there is work and go away when there is not, in three tiers — free, Pro and business.
  • Nothing is left behind in transit. The worker receives the file, returns the result and releases. It does not hold recordings between jobs.

Recognising a speaker again, by consent

Separating voices inside a meeting always works and needs to know nothing about who anyone is. Recognising the same person in a different meeting by their voice is another matter: that is biometric data, so it is off by default. An organization turns it on, and even then only speakers whose consent was recorded are stored and matched. A voiceprint held without consent is deleted after 30 days, and one that goes unused is deleted after 365. There is a plain-language version in speaker diarization, explained.

The summary

The summary is written from the transcript, not from the audio: the language model receives text only, on routing configured not to retain it, and returns a fixed structure — key points, decisions, action items. You can attach an instruction, and ask for a new summary later without transcribing again. Which providers are involved, and on what terms, is set out in the privacy policy rather than implied here.

Why the arrangement matters to you

  • Accuracy: a model built for spoken Hebrew and judged on real meetings, not on clean readings.
  • Privacy: the recording does not go to an external transcription service, and is not used for training.
  • Time: parallel processing with capacity that grows under load.

And the honest part

Automatic transcription still makes mistakes, mostly in names and numbers. That is why the transcript is editable, why every line carries a timestamp back to the audio, and why anything critical should be checked against the recording before you rely on it.

Try it on the hard one

The only test that means anything is your own difficult recording. The free plan covers 7 meetings and 120 minutes per organization per month, with every feature included.

Frequently asked questions

Is the recording sent to another transcription provider?

No. Transcription runs on our own model, on our own infrastructure. The recording is not forwarded to an external speech service.

Does the language model see my audio?

No. The summary is written from the transcript, so the language model receives text and never the recording — and on routing configured not to retain it.

Are my meetings used to train your model?

No. Meetings of registered users are not used to train models. The model improves on material we are licensed to train on, and every candidate is judged on recordings of real meetings before it is allowed near production.

Where is the recording kept?

In your organization's account, until you delete it. Deleting a meeting, a speaker or the whole account removes the stored files with it.

Try it on your own recording

IvreetMeet transcribes and summarizes Hebrew meetings. The free plan covers 7 meetings and 120 minutes per organization each month.

Start free