AI

Local speech recognition in everyday work

4 min read

A sound wave running into a computer and turning into lines of text

Explaining a quote takes two minutes out loud. Typed, it takes a quarter of an hour. Which is why a lot of it never gets written down at all: the short note after a customer visit, the minutes of a meeting, the sentence explaining why you decided the way you did back then. Dictating would be the obvious answer, except that with most services the recording travels to somebody else's server first. In plenty of companies that is where the idea dies, at the latest once customer names start coming up. Local speech recognition on your own machine is therefore not a gadget. It is the difference between "noted" and "I will do that later".

Local speech recognition: the computing power is already on your desk

What has changed over the last few years is not only how well recognition works. It is where it can happen. A current office machine with a decent graphics card transcribes speech faster than a person can produce it. No data centre required, just a model that runs on the device and software that gets out of the way.

We built one for ourselves and use it every day: VoxCast. Press the shortcut, speak, let go, and the text appears in whatever window is active, whether that is the mail client, a ticket system or an editor. Nothing to switch to, nothing to paste.

Two things around that turned out to matter more in daily use than plain dictation:

  • Meetings write themselves. VoxCast records audio from Teams, Zoom or Meet including system sound and transcribes it, so the minutes appear alongside the conversation instead of pulling somebody out of it.
  • Files become text. Recordings and documents turn into clean text by right click, ready to work with.

Recognition happens entirely on the device, in roughly 99 languages. For Windows, VoxCast is a free download.

Three places where it pays off in practice

Dictation is not an end in itself. It earns its keep where something currently does not get written at all, because typing takes longer than the thought did.

After the customer visit. Two minutes of talking in the car, and the CRM holds what was discussed, who owns what and when somebody follows up. The alternative is the note on the passenger seat that survives until Friday.

In support. Anyone who is on the phone and typing at the same time is doing both halfway. The ticket gets written after the call in one go, while the details are still fresh, and it ends up longer than the three keywords it would otherwise have been.

Out in the field. A service technician with oil on his hands is not going to type a report on a tablet. He speaks it, watches the text appear and corrects one line instead of writing three paragraphs.

The pattern is the same every time. Text comes into existence that was not there before. Not text typed faster: documentation that would simply have been missing.

Where your own machine is not enough

Being honest about this belongs to the job. Heavy dialect, several people talking over each other, a bad microphone, and local recognition gets shaky. A cloud service would not do better in that situation either. And once a whole team is involved, you probably do not want the computing load sitting on every laptop. That is what the server version is for: one machine inside your network handles transcription for everybody, and the recording still never leaves the house.

Common questions

What does "local" mean here exactly?

Recording and recognition both happen on the device the person is speaking into. No audio is sent anywhere, no session is created on a service. If somebody switches on the optional AI clean-up of the finished text, that is a deliberate step with their own account behind it.

Is this unproblematic in terms of data protection?

The part that makes the paperwork hard with hosted services, transferring the audio to a processor that often sits outside the EU, does not arise. What remains is the older question of whether you are allowed to record the conversation at all. For meetings you need consent from the participants, no matter where the recognition runs.

And if we would rather use a cloud service?

Then that is a perfectly good decision. Cloud and your own operation are equal routes for us, and we look for the one that suits your data, your team and your budget. It should just be a decision rather than something that happened by default.

How we approach it

VoxCast is a small example of something we build at a larger scale: AI systems that run inside the company itself, from an assistant with access to your own documents to recognition on a server in your own network. The condition never changes. It has to fit into the working day instead of interrupting it.

If a lot gets said in your company and little gets written, tell us where the text is supposed to end up: in the ticket, in the CRM, in the minutes. That is what decides whether a tool is enough or whether an integration belongs with it. Talk to us.

Share:LinkedInX

Hashfox GmbH

Written by the Hashfox team, from live projects, not from the drawing board.

AI that fits the way you work. Let’s talk about your use case.

Explore AI solutions
All insights & news