The Legibility of Live Rewind
As I'm sure most people with even a passing interest will have heard by now, the new Apple Watch has a feature called Live Rewind. Before we get into this properly, let's just see how Apple describe the feature:
Live Rewind, which can help users recall what was just said in conversation. A double press of the Digital Crown shows the previous 15 seconds of a conversation as a text snippet, so the user can catch something they may have missed or are less familiar with. Users can ask Siri about the content of the text or save it to the Siri app to revisit later.
This is what it looks like, apparently:

This example seems to be someone who is hard of hearing catching up what a waiter said in a busy restaurant. As I saw someone mention, usually when a Big Tech company gives the example of helping people with some kind of disability, it's because there's a privacy implication.
Apple are, however, at pains to point out the privacy aspects of how they've implemented what they're calling 'Audio Intelligence':
Audio Intelligence features are private by design, combining hardware, software, and services to protect user privacy in a way that only Apple can. These features do not create or store audio recordings, and raw audio used for processing is completely inaccessible to the operating systems, apps, the user, or Apple. This is because the S11 chip on Apple Watch Series 12 includes Secure Exclave, a dedicated hardware-isolated compartment that processes audio in complete isolation from the rest of the system, then immediately deletes it.
Users are in control and can choose whether to opt in to each Audio Intelligence feature. They can decide exactly where and when Siri Recap takes notes, and can turn Siri Recap on or off at any time, right from Control Center.
By design, Audio Intelligence features do not identify and attribute speakers, to protect the privacy of both the user and those around them. Siri Recap produces a brief, high-level summary of conversations — not a transcript — and is designed to exclude sensitive information like financial information or government-assigned identifiers. Live Rewind plays an audible chime when it is being used, even if Apple Watch is on silent, and shows a full-display animation and microphone indicator so people nearby can hear and see that the feature has been activated.
While I think that people are right to think this is creepy, there are three things which seem reasonable here:
- It's a summary not a transcript
- The identity of speakers isn't included in the summary
- Audio processing happens in a secure enclave and is then deleted
For context: I've started turning on the Zoom transcription by default for my client meetings, because it means I can be more present and can send everyone a summary afterwards. In that case people are entering a meeting space. That's not the same as a chat in a coffee shop.
15 seconds of audio? With an audible chime if it's being used? That's only going to be useful in a very small number of cases, which is why I'm sure that either Apple, or a competitor, will increase that over time. Yes, it's the "thin edge of the wedge" argument, but in this case, I think it's valid.
Dan Hon has some thoughts, too
At issue is this: how ephemeral is speech? When you’re having a conversation with someone, what are the expectations around that speech being recorded, whether the recording is transcription or audio? How long do you expect your words to exist in the world, from day to day? How do those expectations change from context to context, where that context includes both the subject matter, the people you’re conversing with, and more?
People don’t like, I think, in general to be confronted with the exact words they said. When speech is seen to be equivalent with thought (it is, I’d say, the fastest method people have for sharing their thoughts), then what space or grace do you give to someone who is thinking as they’re talking? That’s ultimately the issue here, regardless, unfortunately, of issues like accessibility where live transcription is incredibly useful or critical.
We're talking about a summary of 15 seconds of audio here, an amount of text that it's possible to cram onto a small watch face. But, again, at some point it will become "normal" or at least weird to object to conversations being automatically recorded/summarised. I'm pretty sure that's already the case in some tech hubs, with people using apps like Granola for Apple Watch.
Some people will react to this by saying that people should just be careful with their words. But a text-based reconstruction of what someone said can never fully capture meaning. It doesn't convey the sarcastic tone of voice, the raised eyebrow, or the hand on the shoulder while looking deep into another person's eyes.
Dan, again:
I think people would say that it’s not fair to be judged constantly on your words. For starters, we talk about judging actions, not words. Words are cheap, words can be careless, words can be mere utterances. How long should they stick around? Writing takes effort, perhaps not very much for some people, but I’d argue that writing is also intentional. That’s what’s going on here. The instant that someone’s speech is transformed into another medium -- like text displayed on a screen, or transcribed by a stenographer, or even by someone like me who can type super fast -- then there’s the threat or the fear that it won’t ever go away. ”I didn’t mean that” is a reasonable defence: but then, who else can know whether you meant that or not? The interface between what you think and what you say is known only to the person who made the utterance. The potential for being misunderstood is always there. Transcripts are devoid of context, and I’ve written before that what’s thought of as the technological fix for “not enough context to understand the words as best as possible” is at first glance to gather ever more data, to try to approximate as much context as possible. That’s... not possible. The map, the context you gather, will never be the territory.
For me, it's a question of legibility, and therefore of literacy. By default, there is no permanent, or even semi-permanent, surface to a conversation; it has to be recorded for that to happen. The only way our ancestors could have been presented with their words in a consequential way would be from witnesses in a court of law. That relies on human judgement, which takes in the full-spectrum communication intent of the speaker. Humans can lie and misinterpret, but at least they get the full picture.
We're creating new surfaces by which to be judged. The same goes with glasses and sunglasses that can record other people: we find it creepy because it's just not how we expect the world to work. Returning to Dan Hon, I think he's correct to say:
I don’t think there is a way around this other than for our behavior and expectations about speech to change. ”He said, she said” is a thing people experience, and it’s one that I think we all find difficult to navigate. People go to court over defamation, whether that’s libel (the written one) or slander (the spoken one).
So I don’t think in that other respect that it matters whether Apple is throwing away the recording, whether that’s the audio or the text, because the mere experience of seeing your words repeated back to you is one that people are afraid of. Part of that fear is down to not knowing what’s going to be done with your words. And there’s that phrase: your words. But when they’re written down, what does your words mean anymore? You don’t have control over them anymore, but they’re still attributed to you. You’re still in a way liable for them. They have the ability to become some sort of permanent record.
Again, although Apple commits to deleting the recordings used to process the audio used for the summary, because of their user base they're in effect creating a new sociotechnical system:
Sociotechnical systems [...] is an approach to complex organizational work design that recognizes the interaction between people and technology in workplaces. The term also refers to coherent systems of human relations, technical objects, and cybernetic processes that are inherent to large, complex infrastructures. Social society, and its constituent substructures, qualify as complex sociotechnical systems.
As with AI, there is no public debate as to whether, as a society, we want these things and what the likely effects they are likely to have until they hit the market. The question of whether people do, in fact, actually use them individually is a moot point, as optimising one's own life is different to providing for the flourishing of everyone within society.
For example, most people want security, which is why, especially post 9/11, we have increasing amounts of surveillance. But surveillance limits freedom, and therefore constrains what constitutes a flourishing life for most people within a given population.
What I think critics of Apple's move are saying is that we need a debate about where we're heading with this. Dan Hon again:
Microphones will get smaller. The physical package for processing required to transcribe or record gets smaller. Apple’s latest watches appear to also require an accompany new-enough iPhone to be able to do live rewind, but it’s probably just a matter of time before the transcription can be done solely on a watch.
You don’t know what is going to record you and you don’t know where you’re going to get recorded. You will not know what objects are paying attention to and what objects are not paying attention to you. There don’t need to be people in the room. You will not know who has access to that text. You will know that if it exists, that text could go anywhere, to anyone, and may easily outlive you, even if there’s no present reason to. The cost of there being a permanent record of what you say becomes more and more trivial, and it’s not so much the value in terms of surveillance, but the value in terms of what it could be worth in terms of advertising... but then also it can be stored in case someone wants to buy it in the future.
... and this was all assuming perfect transcription, too.
When I create a "text", whether it's a written one like this one, or an audio or video file, and release it into the world, it is an intentional act. Our conversations with one another are usually not rehearsed, pre-prepared speech acts, but rather a form of "thinking out loud". We cut each other slack, especially when the words don't come out the right way. There is a fundamental ambiguity to our communication, as our words both denote and connote.
These days, it seems that governments have largely abdicated responsibility around regulation to the self-regulation of Big Tech companies. There is tinkering around the edges with populist, but misguided (and ineffective) 'bans' on social media for young people, but while Big Tech companies are hiring philosophers, politicians are mainly interested in the next election cycle.
It's a desperate state of affairs, really. Our political and legal systems aren't fit for purpose, so all we really have is social pushback. That happened with things like Google Glass, with similar technology now smaller and less obvious when baked into Meta's glasses.
So what are we supposed to do when we think that something isn't a good idea? There's no mechanism for having a global debate on this, only putting pressure on companies who decide not to forge ahead with something because it will be reputationally damaging and therefore affect their bottom line.
At the end of the day, everything is about power and who gets to decide how and where we interact with one another. It used to be the state that had that kind of control, backed up with a monopoly on violence. These days, however, that power balance has shifted.
[T]he issue is as it has always been: what does enforcement look like? Who gets to exercise that enforcement? Now we’re back to who has power, and the power differentials that exist in our societies. Who gets to say that no recording happens? Who has that privilege? When is that imbalance recognized? These conversations are recorded for training purposes, but can I record them for consumer rights purposes? Why are consumer rights needed to justify someone recording a conversation? It’s not even clear in some jurisdictions whether you’re allowed to record the actions or speech of a public servant -- police, just so I’m clear -- and even if you’re allowed to record the actions or speech of a public servant, are they going to stop you in practice? Do they in practice exercise the use of force? In which case, what’s a remedy, that the microphones are woven into your clothes, that they’re hidden, that you don’t need to hold a phone up in the first place?
Interesting times, indeed.