notrec technical note 1, October 2026
What counts as a recording? A functional test for always-on AI capture
The notrec maintainers
Abstract. Devices that transcribe, summarise or describe what they sense are increasingly sold as "not recording". We propose a test that does not depend on the word: a capture is a recording if what it keeps lets someone later learn what was said or seen. We hide private facts in a written scene, run nine generic capture pipelines on it, and check which facts can be read back. Six of the nine keep at least one fact. Five of those six come with a "doesn't record" claim that is true as worded.
1. The word is moving
On the watch. TechCrunch reported that the maker says it "does not create or store audio recordings"; a rewind transcript can be saved, and recaps are notes with a title, summary and key points3. On 6 October 2026 a column in The Verge argued that companies building always-on AI devices are redefining "recording" rather than avoiding it1. Its examples: a home camera in development that turns what it sees into text instead of storing video, watch features that transcribe the last fifteen seconds of speech or write a daily recap, and smart glasses pitched as processing what they see without saving footage2.
The point is not that these products are hiding anything. It is that consent rules and indicator lights were built on a binary, a microphone or camera is either recording or it isn't, and a summary sits on neither side of it.
2. A functional definition
Why facts. A summary that says "they talked about plans" is harmless. A summary that keeps a door code is not. Facts make the difference measurable. We call a capture a recording if a person with access to what it keeps can later learn something specific that was said or seen. This is testable. Write a scene, hide facts in it, run the capture, and look for the facts in what stays on file.
3. Method
The control. Every fact has a decoy with one word changed (8203 for 4471). A decoy that matches would mean the index test is guessing. None did. The kitchen scene has three people, 26 things said over four minutes, five things a camera would see, and eight private facts: a flight time, a job not yet announced, a door code, where a spare key is, a wifi password, a surprise party, a name and address on a parcel, and a passport left on a counter. A fact survives in text when all of its key words are present; in raw media when the media comes from the right sensor; and in a searchable index when a probe for its words matches a stored vector.
The nine pipelines are generic stand-ins, not models of any product: live captions, sound events, a talk-time counter, a wake-word assistant, a fifteen-second rewind saved at two taps, a six-sentence daily recap, a camera that writes descriptions, a searchable memory of hashed word vectors, and an audio recorder.
4. Results
Wake words. "Only listens after the wake word" is true. The person who speaks next is not the one who said it. Captions, sound events and talk-time counts keep nothing that brings a fact back. Everything else does (Table 1, computed in your browser). The daily recap keeps four facts because a good summary keeps what matters, and what matters is names, numbers and dates. The wake-word assistant keeps one: the door code, said by someone else seconds after the owner spoke to it. The searchable memory keeps no sentence at all and still gives away six facts to a simple probe.
| Pipeline | Claim | Class | Facts back | On file |
|---|---|---|---|---|
| Computing… | ||||
5. Classes and the light
Why red for a gist. Because of Section 4. A recap is a gist and kept four facts. We sort captures into five classes by what stays on file: nothing (R0), signals such as times and counts (R1), gist such as a summary or index (R2), content as text (R3), and raw media (R4). The light follows the class: white while a device listens and keeps nothing, amber while it keeps signals, and red, with everyone present told, once it keeps a gist or more.
6. A gate instead of a promise
The library's consent gate keeps words only while everyone present has said yes. Replaying the kitchen scene, one guest saying no brings the kept facts from 8 to 0; with everyone saying yes and numbers and pass-phrases scrubbed, 5 of 8 remain.
7. Limits
The scenes are written, so pipelines work on text; real speech recognition makes mistakes and a language-model summary rephrases, but both keep names, numbers and dates. The vector probe is a lower bound on what a real embedding index gives away. This is a test and a design rule, not legal advice.
References
- "We can't just change the definition of 'recording'," The Verge, October 2026. theverge.com
- "Apple, Google redraw 'recording' for always-on AI gadgets," AI Weekly, 6 October 2026. aiweekly.co
- "Apple Watch's new AI features are normalizing the idea that technology is always listening," TechCrunch, 9 September 2026. techcrunch.com