AI Regulation Tracker / Enforcement
Korea fined Apple's distribution arm over Siri transcripts, not just the recordings
The Personal Information Protection Commission treated the text derived from voice-assistant audio as a separate processing stream needing its own lawful basis. Fixing consent for the audio did not fix the transcript.
What did the PIPC actually decide?
Two matters, one sitting, one combined number. The release states that the Commission imposed administrative fines totalling KRW 10.558 billion on TikTok Pte. Ltd. and two Apple affiliates, with corrective and publication orders, for violating the Personal Information Protection Act and the former Network Act.
The TikTok side is the larger figure and the more familiar fact pattern. PIPC fined TikTok Pte. Ltd. KRW 10.306 billion over third-party behavioural data collected through TikTok Pixel, the Events SDK and the Events API. The scale is the point: 9.45 million Korean users, roughly 71,000 Korean sites and apps, no lawful basis, no valid separate consent, and an unlawful cross-border transfer finding on top.
The Apple side carries a much smaller number and a much larger idea. Two corporate entities, treated differently. Apple Distribution International was fined KRW 252 million. Apple Services Pte. Ltd. received a corrective order. Summarising that as a fine against "Apple" loses the part that matters for group-structure analysis: the regulator identified which affiliate did what and matched the remedy to it.
Why is the transcript holding the part that travels?
Here is the operative finding, in PIPC's own framing. Until August 2019, when users used the voice assistant service Siri, voice recordings and transcripts were collected and used for purposes including improving voice-recognition functions and search results, without separate consent from users. Then the novel half: although from October 2019 separate consent was obtained for using voice recordings for service improvement, the transcripts continued to be used for service improvement without a separate lawful basis being established.
Consent was fixed for the audio. The text produced from the audio kept running on nothing.The finding, in plain terms
Read that again if you run a speech pipeline. The regulator did not treat the transcript as a format of the recording. It treated it as a distinct processing stream with its own consent requirement, and found that repairing the upstream consent did not carry downstream.
That is a structural holding, and structural holdings are the ones that get borrowed. It reaches any pipeline where a raw input is captured under one basis and then transformed into a durable secondary artefact: audio to transcript, transcript to summary, clinical conversation to structured note.
Who is affected, and which entity got what?
The direct exposure sits with foreign platform and AI providers serving Korean users: voice and speech-recognition product teams, engineers running human-review or transcript pipelines for model improvement, adtech teams deploying pixels and SDKs, privacy counsel at non-Korean providers, and whoever owns cross-border transfer paperwork.
The indirect exposure is wider, and it is where the US professional should be paying attention. Ambient AI scribes are now standard equipment in American clinical and legal settings, and the consent architecture in most of those deployments is built around the recording. The patient or client is asked to consent to being recorded. The transcript, the structured note and the model-improvement loop get treated as downstream plumbing covered by that same consent, when they are described at all.
Korea has now issued a decision that does not accept that framing. It binds nobody in the United States. It is still the kind of reasoning that shows up in the next regulator's analysis, because it is cheap to adopt and hard to argue against on the text.
How does this compare with the EU and California on derived data?
The comparison below is structural, not a claim that identical rulings exist elsewhere. It asks whether each framework's text supports treating a derived transcript as needing its own basis.
| Regime | Does the derived text fall within the protected category? | Is a fresh basis needed for a new purpose? | What this decision adds |
|---|---|---|---|
| South Korea, PIPA and the former Network Act | Yes on these facts. The transcript was assessed on its own footing, not as an attribute of the recording. | Yes. Separate consent or another lawful basis is required for use in service improvement. | Applied and enforced. Consent repaired at the recording layer in October 2019 did not cure the transcript layer. |
| European Union, GDPR | Text supports it. Article 4(1) defines personal data as any information relating to an identified or identifiable natural person, regardless of how it was produced. | Text supports it. Article 6(1) requires a lawful basis, and Article 5(1)(b) limits further processing to compatible purposes. | A worked example of a regulator reaching the derived layer separately rather than folding it into the source. |
| California, CCPA as amended by the CPRA | Text supports it. The definition of personal information at Civil Code section 1798.140(v)(1)(K) expressly reaches inferences drawn from other personal information. | Partly. Use is limited to purposes disclosed at collection or compatible with them, rather than requiring consent for each derived use. | Shows a regulator splitting a pipeline into layers, the same analytical move a California enforcement theory would need. |
The pattern is the same in all three, and it is not new law anywhere. What Korea supplied is the demonstration. A regulator walked a pipeline, found two layers, and held that fixing one did not fix the other.
What should a compliance officer do about speech-to-text pipelines?
Start by drawing the pipeline honestly. Most consent records describe a product, not a data flow. If your notice says the user consents to voice recording for service improvement, and your architecture keeps a transcript after the audio is deleted, you have two artefacts and one authorisation. Write down every artefact the pipeline creates and how long each one lives.
Then check whether retention matches. Speech systems commonly pair a short audio window with an indefinite transcript store, on the theory that the transcript is less sensitive. Korea's decision runs the other way.
Third, look at your remediation history. The Apple finding has teeth because a fix made in October 2019 was found not to have covered the whole problem. Remediations get scoped to whatever caused the complaint. Go back and confirm what yours actually covered.
Fourth, treat human review as its own question. Transcripts used for model improvement are usually what human reviewers see.
None of this replaces counsel's judgment on your own facts. It is the evidence-gathering that has to happen first.
What this decision does not say
It does not ban AI training on voice data, and it announces no general rule about generative-AI training corpora. The release addresses consent and lawful basis for one voice-assistant improvement pipeline and one behavioural-advertising pipeline, plus cross-border transfer. Reading it as a Korean position on foundation-model training data goes past the text.
It also settles nothing outside Korea. This is an administrative sanction under national statute, and the parties retain whatever review rights Korean administrative law gives them. The value to a US reader is the analytical structure.
Frequently asked questions
How much did Korea's PIPC fine TikTok and Apple?
PIPC resolved on 22 July 2026 to impose fines totalling KRW 10.558 billion: KRW 10.306 billion on TikTok Pte. Ltd. and KRW 252 million on Apple Distribution International. Corrective and publication orders accompanied them, including a corrective order to Apple Services Pte. Ltd.
What was the finding about Siri transcripts?
PIPC found that until August 2019, Siri voice recordings and their transcripts were collected and used for purposes including improving voice recognition and search results without separate user consent. It further found that although separate consent for the recordings was obtained from October 2019, the transcripts were still used for service improvement without a separate lawful basis.
Did PIPC ban AI training on voice data?
No. The resolution addresses consent and lawful basis for a specific voice-assistant improvement pipeline and a specific behavioural-advertising pipeline, plus cross-border transfer. It does not state a general rule about AI training corpora.
What was TikTok fined for?
PIPC fined TikTok Pte. Ltd. KRW 10.306 billion for collecting third-party behavioural data through TikTok Pixel, the Events SDK and the Events API from 9.45 million Korean users across roughly 71,000 Korean sites and apps, without a lawful basis or valid separate consent, and for unlawful cross-border transfer.
Last verified: July 28, 2026