Workspace Serenity: Tech and Tea on a Wooden Desk

How async standups work when the update is a voice note

by admin

An async standup replaces the daily synchronous meeting with written updates that people post and read on their own schedule, and for teams spread across more than about four hours of timezone difference it is usually the only version that works at all. The theory is sound. Nobody wakes at five to say three sentences.

Then someone on the team starts posting voice notes instead of typing, because it is faster and because they are walking, and within a fortnight half the team has switched. The standup still happens. Reading it becomes impossible.

How does an async standup work?

An async standup works by having each team member post a short update at a time that suits them, covering what they completed, what they are working on next, and what is blocking them. The updates go into a shared channel or a tracker thread rather than being spoken aloud in a meeting, and colleagues read them when they start their own day. The mechanism only functions if the updates are genuinely readable at speed, because the entire time saving comes from the fact that reading eight updates takes two minutes while listening to eight people takes fifteen.

That is the design constraint the whole format rests on, and voice notes break it.

Cheap to produce, expensive to consume

This is the asymmetry that makes voice updates spread and then fail.

Producing a voice note costs the speaker almost nothing. Two minutes of talking, no editing, no formatting, no self consciousness about phrasing. Producing an equivalent written update costs perhaps four minutes and requires being at a keyboard. From the producer’s side, voice wins clearly.

Consuming is the reverse and it is not close. A written update is scanned in fifteen seconds and skipped entirely if it does not concern you. A voice note cannot be scanned. It has to be played, in order, at the speed the speaker chose, and the listener has no way to know whether the relevant part is at second ten or second ninety without listening to both.

Nielsen’s foundational finding on reading behaviour was that 79 percent of users scanned any new page rather than reading word by word. Scanning is the default mode for exactly this kind of routine, low stakes, high volume information. Audio does not permit it.

So a team of eight generates sixteen minutes of audio per day, and each person is nominally expected to consume all of it. Nobody does. What happens instead is that people listen to the two colleagues whose work touches theirs and ignore the rest, at which point the standup has stopped being a team wide sync and become a set of private channels with extra ceremony.

What the standup is actually for

Worth being precise here, because the answer determines what the format has to support.

A standup exists to surface blockers early, to prevent duplicated work, and to give the team a shared picture of where things stand. Of those three, only the first has an urgent time dimension. The second and third are served just as well by a record that can be read later as by one that is consumed today.

So a standup is closer to a documentation event that has been performed live for historical reasons than to a communication event. Once that is clear, the criteria for a good async format follow directly. It has to be searchable, linkable from an issue, and skimmable by someone catching up after a week off.

Audio is none of the three. A voice note cannot be pasted into a ticket as context. It cannot be found six weeks later when someone asks when the migration actually started. It cannot be skimmed by the person returning from leave, who now has forty voice notes and will simply not listen to them.

Keep the voice, add the text

Banning voice notes is the wrong move. The fix is to stop treating the audio as the artefact.

Let people record. Speaking is faster, it captures more nuance about uncertainty than a bulleted list does, and it is genuinely more accessible for some team members. Then transcribe, so that what lands in the channel is text, with the audio attached for anyone who wants it.

This is a small pipeline and it can be entirely automatic. A batch of short recordings goes in and a set of summaries with extracted action items comes out. Tools built for this, https://vomo.ai among them, will take a group of recordings, produce transcripts with speaker labels, generate a summary per recording, and pull out the items that look like commitments or blockers. Vomo exports as text, Markdown, DOCX or PDF, and Markdown is the one that matters here because it drops straight into most trackers and wikis without reformatting.

None of this is automation for its own sake. The artefact the team keeps should be the one that can be searched, and the artefact the individual produces should be the one that is fastest for them. Those do not have to be the same file.

The accessibility argument, which is not optional

There is a version of this argument that treats transcripts as a nice extra. The stronger version is the one that matters here.

The W3C’s guidance on transcripts for audio content is explicit that basic transcripts serve people who are deaf or hard of hearing, people who have difficulty processing auditory information, and people who simply process text better than audio. A team standup that exists only as audio is a standup that some colleagues cannot participate in on equal terms.

For a public facing product this is a compliance conversation. For an internal daily ritual it is a straightforward question of whether everyone on the team can actually read the team’s own updates, and the answer for an audio only format is no.

Three formats measured against what a standup has to do

Setting the options side by side makes the trade clear, and it also shows why the hybrid is not a compromise but a strict improvement.

  Live standup Voice notes only Voice recorded, text published
Cost to the person updating Low, but at a fixed hour Lowest Lowest
Cost to the person catching up Must attend or miss it Roughly one minute per teammate Roughly ten seconds per teammate
Searchable six weeks later No No Yes
Linkable from an issue No No Yes
Usable by someone who cannot listen Partly No Yes
Works across eight timezones No Yes Yes

The middle column is where most teams end up by accident. It is better than the left column on exactly one axis, timezone tolerance, and worse or equal on everything else, which is a poor trade for a daily ritual.

The right hand column costs one automated step. Everything else about how the team behaves stays the same, which is the reason it tends to stick when process changes usually do not: nobody is being asked to work differently, only to let the recording become text before it reaches anyone.

Where async standups genuinely fail

Two failure modes, and neither is fixed by transcription, so they are worth naming separately.

The first is that async standups do not surface urgent blockers fast enough. If someone is blocked at nine in the morning and the person who can unblock them reads the update at four in the afternoon, seven hours were lost that a live meeting would have saved. Changing the standup format does nothing here. What is needed is a separate escalation path for blockers, used immediately rather than saved for the daily update. Teams that put blockers in the standup are using the wrong channel.

The second is that async updates drift into status theatre. Written updates are performative in a way spoken ones are not, because they persist and are read by managers. The reliable symptom is updates that grow longer over time while containing less. The fix is a hard format limit, three lines or thereabouts, enforced by convention.

A format that survives

For teams that want to try this, the shape that works is unremarkable and specific.

Record for no more than ninety seconds. Longer updates are a signal that something needs a document or a call, not a standup slot.

Say your name at the start. It makes speaker attribution reliable and costs one second.

Cover three things in a fixed order: what moved since yesterday, what is next, what is blocked. Fixed order matters more for reading than for speaking, because a reader scanning eight updates is looking at the same position in each.

Post the text as the primary artefact and attach the audio. Not the other way round.

And link out to issue numbers explicitly, by saying the number aloud. A transcript containing a ticket reference becomes searchable from that ticket, which is the point at which the standup archive stops being a log and starts being useful project history.

Nobody has to give up the thing that made voice notes spread in the first place, which was simply that talking is faster than typing. The change is that what is fast for the person writing stops being slow for the seven people reading.

Related articles

E-commerce Platform
How To Safely Change Your eCommerce Platform

Migrating from one eCommerce system provider to another is probably the most stressful decisions an eCommerce retailer has to make.…

How AI Agent Builders Can Optimize WordPress Performance Workflows
How AI Agent Builders Can Optimize WordPress Performance Workflows

In today’s fast-paced digital environment, website performance is no longer optional—it’s critical. A slow-loading WordPress site can lead to higher…

Elementor
20 Tips for Creating a Visually Appealing Website with Elementor

20 Tips for Creating a Visually Appealing Website with Elementor First impressions matter, and when it comes to your website,…

Ready to get started?

Purchase your first license and see why 1,500,000+ websites globally around the world trust us.