The Terms You Agreed To
Biometrics and the modern photographer
A DC headshot photographer named Moshe Zusman posted earlier this summer about his wife using one of the AI headshot services she’d been seeing advertised. She works in marketing, so this was research, apparently. She uploaded a few phone photos, got back something polished enough that she liked the lighting and her hair, and then told her husband about it.
He asked whether she’d read the terms of service first. She hadn’t.
I can’t remember the last time I read every word of a terms of service agreement before clicking whatever button stood between me and the software I was trying to use. Zusman pointed out that she’d handed photographs of her face to a platform she knew little about, with no idea whether they might be used to train models or end up in a facial recognition database. She deleted the account that night.
I kept coming back to the privacy part, partly because I’d never looked at my own software that way.
What the file becomes
My first concern was unfounded. Photographs themselves are not treated as biometric data.
Illinois’ Biometric Information Privacy Act, probably the best-known biometric privacy law in the country, explicitly excludes photographs from its definition of a biometric identifier. It includes scans of face geometry. It also includes voiceprints.
The photograph or recording is the source material. Software can derive something else from it.
Note taking has never been a skill I excelled at, so about a year ago I subscribed to Otter.ai. Otter is transcription software that helps me improve my client experience. With permission, as I talk to a client, Otter records the conversation so I don’t have to interrupt to take notes or ask someone to repeat themselves because I missed a word. It’s hard to overstate how much better it is to have a searchable transcript instead of whatever notes I managed to scribble while also trying to pay attention to the person in front of me.
The terms describe another process. Otter calls the result Speaker Identification Information. Its Data Processing Agreement is less delicate about the terminology: the service may process voiceprints so it can recognize users and automatically tag their names in a transcript.
That isn’t something I need it to do, partly because I was in the meeting and partly because Otter doesn’t even do a great job with that task. I can’t tell you how many meetings I’ve been in with a few people and ended up with a transcript that assigns twice the number of speakers as were actually in the meeting.
A transcription service can hear that one person stopped talking and another one started and label them Speaker 1 and Speaker 2. It doesn’t need to know either person’s identity, and it doesn’t need to remember the voice after the meeting ends.
Otter can remember enough about a person’s voice to recognize that person later. I had agreed to that. The client whose voice I recorded hadn’t agreed to Otter’s terms at all.
I had been asking the wrong question (it seems I do that a lot)
My first concern was model training, mostly because AI companies have trained the rest of us to look for that language now.
Otter does use some customer material for that purpose. Its current privacy documentation says it can train its proprietary AI using de-identified audio recordings and transcriptions. The transcripts may still contain personal information, something Otter itself acknowledges. Of course they can. Otter doesn’t control what people say during the recording.
The voiceprint appears to be a separate operation. Otter’s public documentation doesn’t say that it trains models on the voiceprint itself. It uses Speaker Identification Information to recognize speakers.
“They train on your data” turned out to be too blunt a description of what was happening. The same product makes several decisions about the same conversation: what gets recorded, what gets retained, what gets derived from the recording, whether that derived information identifies someone later, and whether any of the original material enters a model-training process.
Otter’s Enterprise controls made that easier to see. Enterprise customers can disable speaker learning. When they do, Otter can still separate speakers in a transcript, but it stops creating or updating the persistent voice information used to identify them later. Otter describes the setting as useful for organizations with biometric, privacy, or compliance requirements. The switch lives in Enterprise.
I had also been trying to understand how all of this squared with Otter advertising HIPAA compliance. The answer is less contradictory than it looked at first. HIPAA-compliant Otter isn’t the ordinary service with a privacy badge attached. It requires an Enterprise account and a signed Business Associate Agreement. Enterprise customers are opted out of global AI model training by default, and they have privacy controls that ordinary accounts don’t. That resolved my HIPAA question.
What I was left with was the feeling that if it costs more to protect my, and my clients’ biometric privacy, then it is that data that covers the rest of the cost. We’ve seen that playbook before.
Three transcription tools
I had already been looking at Fireflies.ai as a possible replacement for Otter, so I read those documents next.
Fireflies also processes characteristics of a person’s voice. Its Privacy Policy says service providers may derive voice information so the software can distinguish one participant from another, and it acknowledges that some jurisdictions may treat that information as biometric data.
Fireflies draws the line earlier than Otter does. Its policy says the voice information isn’t used to identify or authenticate an individual. Fireflies says it doesn’t receive or process that biometric representation on its own servers, and it publishes a destruction schedule for biometric information where Illinois law applies.
The service needs to know that Speaker 1 and Speaker 2 aren’t the same person. It doesn’t need Speaker 2 to still be Michael next Thursday.
The Terms deserve a little more attention than that comparison alone suggests. Section 5(c) says Fireflies won’t use User Content to train, retrain, fine-tune, or otherwise improve generative AI models, but it also reserves the right to derive usage, statistical, and other data from that content for internal business purposes, including service improvement. Its Data Processing Addendum makes a similar allowance for de-identified data. I don’t read that as nothing.
The difference for the question I was trying to answer is that Fireflies gets much more specific about the voice data itself. Its service providers use those characteristics to tell speakers apart, not to identify or authenticate them, and Fireflies says that biometric representation never reaches its own servers.
Fireflies also says meeting content isn’t used to train AI models and that the vendors processing that material operate under restrictions against using it for their own model training. That was considerably clearer than I expected from the first pass through its Terms.
I use Zoom too, which supplied a third version of the same decision. Zoom’s terms explicitly exclude the audio, video, chat, screen sharing, attachments, and similar customer communications in a meeting from training Zoom’s or third-party AI models. Zoom can create a voiceprint. The interesting part is when.
Its ordinary transcription functions can distinguish speakers without permanently knowing who they are. Features that actually need a persistent voice identity, such as personalized audio isolation or automatic smart name tags, require the person whose voice is being learned to enable or enroll in that feature.
I kept expecting the three services to converge once I read far enough into the legal language. They didn’t. They all need audio to transcribe a conversation. They don’t all make the same decisions about what else to extract from it.
The consent chain
Photographers aren’t new to handling other people’s data.
I have redundant systems for image files because losing a client’s photographs would be unacceptable. I use payment systems designed to keep credit card information out of my hands as much as possible. Galleries have access controls. Contracts and client records live in systems chosen partly because those systems are supposed to be safer places for them than a random folder on my computer.
I’d put transcription in a different mental category. It was a convenience and it was customer service. A design consultation can cover where the finished photographs will hang, what size a client is considering, details about the session, names, dates, scheduling, pets, family members, sometimes things that don’t matter to the photography at all but come up because two people are having a conversation. I started using transcription because it let me listen instead of dividing my attention between the client and my notes.
Once I add Otter to that conversation, there are at least two consent questions. One is easy to see. Does the client know I’m recording? The answer is always yes, because I ask for permission before we start. The second took me longer. Does the client know what the service receiving that recording does after it gets there? Here, I dropped the ball.
Recording-consent law varies by state. Biometric law varies much less, because most states do not have biometric privacy laws on the books today. Illinois requires notice and written consent before a private company collects covered biometric identifiers and gives individuals a private right of action when companies violate the statute. Other states regulate biometrics differently, often through their Attorneys General rather than private lawsuits. Arizona doesn’t have a general BIPA-style biometric privacy statute. Lawmakers introduced another biometric privacy bill in 2026, so that’s subject to change.
I don’t think photographers need to become privacy lawyers. But I also think “my state doesn’t prohibit this” is a remarkably low standard for deciding what I’m comfortable doing with somebody else’s information.
What changed
I started this research because I read about someone I don’t know that uploaded photographs to an AI headshot service on LinkedIn. I ended up looking much harder at software I already use.
Otter isn’t secretly doing something it never disclosed. The voiceprint language is there. So is the model-training language. Fireflies has its language. Zoom has its language. I hadn’t read enough of it.
I use transcription because I want to pay better attention to clients. That part hasn’t changed. What changed is whether I’m willing to hand their conversations to a service that can create a persistent representation of their voices when another service can do the job without making the same choice.
My Otter subscription will expire next week and it will not be renewed.


