Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII) Review
0. Introduction
한 줄 요약: 이 position paper는 deepfake research가 authenticity detection과 viewer deception에 과도하게 집중하면서 AI-generated non-consensual intimate imagery, AIG-NCII의 핵심인 consent, identity misuse, dignity harm를 놓치고 있다고 비판하고, prevention, access control, survivor-centered evaluation으로 research agenda를 재정렬할 것을 제안한다.
이 논문을 지금 볼 가치가 있는 이유는 다음과 같음.
- Deepfake detection이 기술적으로 성공해도 피해 image가 계속 유통된다면 실제 harm reduction으로 이어지지 않을 수 있다는 문제를 정면으로 다룬다.
- Synthetic or authentic이라는 축과 consensual or non-consensual이라는 축을 분리해 threat model의 기본 단위를 다시 정의한다.
- Detection, watermarking, provenance가 어떤 harm에는 유효하지만 AIG-NCII에서는 오히려 새로운 risk를 만들 수 있음을 설명한다.
- 기존 AI/ML publication landscape가 AIG-NCII-specific implementation을 얼마나 다루는지 structured review로 점검한다.
- Technical intervention을 benchmark score가 아니라 dignity, autonomy, exposure reduction, misuse friction으로 평가해야 한다고 주장한다.
이 글은 새로운 detector를 제안하는 논문이 아니다. Deepfake research가 무엇을 target으로 삼아왔고, 그 target이 실제 피해자의 문제와 맞는지를 묻는 position paper다.
Deepfake 분야의 전형적인 질문은 다음과 같다.
이 image or video가 real인가 synthetic인가?
AIG-NCII에서 더 먼저 물어야 할 질문은 다르다.
이 사람의 likeness와 intimate representation이 consent 없이 생성, 변형, 소유, 유통되고 있는가?
두 질문은 겹칠 수 있지만 동일하지 않다. Authenticity detector는 synthetic content를 잘 잡아도 consent를 판단하지 못한다. 반대로 real image라고 판정된 content도 non-consensual일 수 있다.
1. Problem Setting
1-1. Problem definition
AIG-NCII는 AI system으로 생성하거나 변형한 non-consensual intimate imagery를 뜻한다. 여기서 핵심 harm는 content가 사실인지 거짓인지에만 있지 않다.
논문은 harm를 두 범주로 나눈다.
- Epistemic harm
- Viewer가 synthetic content를 authentic record로 오해한다.
- False information이 reputation, trust, public discourse에 영향을 준다.
- Detection and provenance가 직접 대응하기 좋은 영역이다.
- Dignity harm
- Subject의 identity와 likeness가 consent 없이 intimate context에 사용된다.
- Content가 명백히 synthetic하더라도 autonomy, dignity, privacy, safety가 침해된다.
- Authenticity label만으로는 harm가 사라지지 않는다.
이 구분을 두 축으로 정리하면 다음과 같다.
| Authenticity | Consent | Example category | Main concern |
|---|---|---|---|
| Authentic | Consensual | Voluntarily shared authentic content | Privacy and distribution context |
| Authentic | Non-consensual | Traditional NCII | Dignity, privacy, coercion, exposure |
| Synthetic | Consensual | Consensual generative or fictional content | Disclosure depending on context |
| Synthetic | Non-consensual | AIG-NCII | Identity misuse, dignity, autonomy, exposure |
Artificiality와 consent는 orthogonal하다. Synthetic이라는 사실은 자동으로 harmful을 뜻하지 않고, authentic이라는 사실은 safe를 뜻하지 않는다.
1-2. Why previous approaches are insufficient
AI/ML deepfake research는 주로 세 가지 technical paradigm에 집중해왔다.
- Detection
- Pixel, frequency, physiological, temporal artifact로 synthetic content를 분류한다.
- Goal은 real vs fake classification accuracy다.
- Provenance
- Content origin, editing history, signing chain을 추적한다.
- Goal은 source and transformation history의 verifiability다.
- Watermarking
- Generator output에 detectable signal을 삽입한다.
- Goal은 model-generated content의 later identification이다.
이 방법들은 misinformation, fraud, evidence integrity에는 중요하다. 그러나 AIG-NCII에서는 다음 gap이 생긴다.
Authenticity is not consent
Detector가 image를 synthetic으로 정확히 label해도 subject가 consent하지 않았다는 사실을 자동으로 판정하지 못한다. Harmful content에 AI-generated label을 붙인 채 계속 공개하면 viewer deception은 줄어도 exposure and dignity harm는 유지될 수 있다.
Public label can become insufficient remediation
Platform가 harmful image를 제거하지 않고 authenticity warning만 붙이는 경우, technical system은 moderation action을 대신하는 표식이 될 수 있다. AIG-NCII에서는 visibility reduction, removal, distribution friction이 더 직접적인 outcome일 수 있다.
Authenticity tools can be repurposed
A high-confidence detector or provenance system이 harmful collection을 정리하거나 특정 content의 origin을 확인하는 데 악용될 가능성도 있다. Tool의 intended use가 verification이어도 attacker에게 sorting or validation capability를 줄 수 있다.
Definitive authenticity claims can harm traditional NCII victims
Traditional NCII victim은 disputed authenticity or uncertainty를 protective ambiguity로 활용할 수 있다. System이 content를 definitive authentic으로 label하면 plausible deniability를 제거해 secondary harm를 만들 수 있다.
따라서 research objective는 단순한 detection accuracy를 넘어 어떤 action과 harm outcome으로 연결되는지를 포함해야 한다.
2. Core Idea
2-1. Main contribution
논문의 핵심 contribution은 세 가지다.
- Misalignment diagnosis
- Current deepfake research는 viewer-centric epistemic harm를 중심으로 설계되어 있다.
- AIG-NCII의 subject-centric dignity harm는 threat model, dataset, metric에서 주변화된다.
- Publication landscape analysis
- High-visibility AI/ML research에서 NCII and AIG-NCII가 실제로 얼마나 다뤄지는지 structured screening을 수행한다.
- Mention과 technical implementation을 구분한다.
- Research agenda
- Detection-centric defense를 넘어 identity protection, access control, preventive intervention, harm-aligned metrics, domain partnership, survivor plurality를 제안한다.
이 position의 핵심 문장은 다음처럼 압축할 수 있다.
Deepfake system이 fake를 잘 찾는 것과 AIG-NCII harm를 잘 줄이는 것은 서로 다른 optimization problem이다.
2-2. Design intuition
논문은 technical pipeline을 harm pathway와 연결해야 한다고 본다.
기존 pipeline은 대략 다음과 같다.
\[\text{content} \rightarrow \text{authenticity classifier} \rightarrow \text{real or fake label}\]AIG-NCII-oriented pipeline은 더 많은 decision을 포함해야 한다.
\[\text{content and context} \rightarrow \text{identity and consent risk assessment} \rightarrow \text{access or moderation action} \rightarrow \text{exposure reduction and support}\]여기서 technical metric도 바뀐다.
- Detector AUC만으로는 부족하다.
- Harmful-content exposure time을 줄였는가.
- False positive가 consensual creator or marginalized community에 어떤 cost를 주는가.
- Identity misuse를 creation stage에서 어렵게 만들었는가.
- Victim report가 신속한 removal로 이어지는가.
- Defense가 adaptive attacker에게 reverse-engineering signal을 주는가.
즉 model output correctness가 아니라 intervention consequence가 evaluation target이 된다.
3. Architecture / Method
3-1. Overview
| Item | Description |
|---|---|
| Paper type | Position paper plus publication landscape analysis |
| Main target | AI-generated non-consensual intimate imagery, AIG-NCII |
| Core distinction | Epistemic harm vs dignity harm |
| Key axes | Artificiality and consent |
| Reviewed paradigms | Detection, provenance, watermarking |
| Main diagnosis | Authenticity-centric research is not sufficient for consent-centric harm |
| Proposed direction | Prevention, identity protection, gated access, harm-aligned evaluation |
| Governance principle | Survivor-centered and domain-partnered research |
3-2. Module breakdown
1) Harm taxonomy
논문은 deepfake harm를 viewer와 subject 관점으로 분리한다.
- Viewer-centric question
- Can a viewer tell whether this content is authentic?
- Can misinformation or fraud be prevented?
- Subject-centric question
- Was this person’s identity used with consent?
- Can creation and distribution be prevented or interrupted?
- Does the response restore control and reduce exposure?
이 taxonomy는 epistemic harm가 덜 중요하다는 주장이 아니다. 두 harm가 다른 mechanism and remedy를 요구한다는 주장이다.
2) Authenticity-consent matrix
논문의 가장 중요한 conceptual tool은 artificiality and consent를 separate axes로 두는 것이다.
- Authenticity detector는 horizontal axis만 추정한다.
- AIG-NCII policy는 vertical consent axis를 함께 다뤄야 한다.
- Consent는 image pixel만으로 안정적으로 infer할 수 없는 context-dependent property다.
따라서 pure content classifier 하나로 complete solution을 만들 수 없다. Identity claim, source, uploader relationship, report, platform context, distribution intent 같은 external information이 필요하다.
3) Research landscape screening
논문은 Google Scholar query로 initial literature pool을 구성한다.
- Initial search result: 965 papers
- Venue and citation filtering: 379 papers
- Top-ranked 100 papers: manual screening
- Final relevant deepfake research set: 39 papers
이 39편을 AIG-NCII relation으로 분류한 결과는 다음과 같다.
| Category | Count | Meaning |
|---|---|---|
| No mention | 34 | NCII or AIG-NCII를 다루지 않음 |
| Mention only | 5 | Motivation or discussion에서 언급하지만 technical method는 일반 deepfake task에 머묾 |
| AIG-NCII-specific implementation | 0 | Threat model, dataset, metric, intervention이 AIG-NCII에 맞춰진 technical work 없음 |
이 결과는 전체 분야의 exhaustive census가 아니다. 저자들이 선택한 search and ranking protocol 안에서 high-visibility research가 무엇을 foreground하는지 보여주는 landscape signal이다.
4) Detection, provenance, watermarking critique
논문은 세 method를 폐기하자고 주장하지 않는다. 각각의 utility와 mismatch를 분리한다.
| Method | Useful for | AIG-NCII mismatch |
|---|---|---|
| Detection | Synthetic-content triage, forensic support | Consent를 판단하지 못하고 label-only response로 끝날 수 있음 |
| Provenance | Origin and edit history | Source trace가 removal or dignity restoration을 보장하지 않음 |
| Watermarking | Cooperative generator output identification | Open-source, removed watermark, non-cooperative attacker에 취약 |
핵심은 method가 어떤 downstream action에 연결되는지다. Detection은 backend moderation triage로 쓰일 때 유용할 수 있지만, public authenticity label만 제공하는 것은 충분하지 않을 수 있다.
5) Research recommendations
논문은 아홉 방향을 제안한다.
- Decouple epistemic and dignity harms
- Detection score와 harm reduction을 별도 objective로 둔다.
- Elevate dignity in threat models
- Publicly available image를 consent의 proxy로 취급하지 않는다.
- Identity preservation and minimization을 포함한다.
- Gate high-risk capabilities
- Identity-conditioned generation, inpainting, face-specific adaptation asset의 access를 risk-sensitive하게 관리한다.
- Invest in proactive prevention
- Adversarial immunization or protective perturbation처럼 creation 단계의 misuse friction을 높인다.
- Build safety-aligned metrics
- Detection accuracy뿐 아니라 exposure, removal, false accusation, victim burden을 측정한다.
- Integrate AIG-NCII into AI safety
- Cooperative model or visible watermark만 가정하지 않고 non-cooperative operator를 threat model에 넣는다.
- Establish domain partnerships and guardrails
- Legal, policy, victim-support, platform moderation expertise와 공동 설계한다.
- Respect survivor plurality
- 피해 경험과 원하는 remedy가 단일하지 않음을 반영한다.
- Include social intervention
- Technical defense만으로 해결하지 않고 education, reporting support, platform policy, law enforcement process를 함께 본다.
4. Training / Data / Recipe
4-1. Data
이 논문은 model training dataset을 제안하지 않는다. 대신 publication landscape와 conceptual cases를 evidence로 사용한다.
Literature pipeline은 다음 data source를 포함한다.
- Google Scholar search result
- Venue and citation metadata
- Top-ranked paper title, abstract, and full-text screening
- AIG-NCII mention and implementation status
Screening의 중요한 distinction은 mention과 method다. Introduction에서 harm를 motivation으로 언급했더라도 다음이 없다면 AIG-NCII-specific implementation으로 보지 않는다.
- Consent-aware threat model
- Identity or subject-centric dataset design
- AIG-NCII-specific metric
- Prevention or removal intervention
- Survivor or domain-expert evaluation
4-2. Training strategy
Training recipe 대신 research-design recipe를 정리하면 다음과 같다.
- Harm definition before model design
- 누구의 어떤 harm를 줄일 것인지 명시한다.
- Threat actor and deployment context
- Cooperative platform, closed model, open-source model, malicious operator를 구분한다.
- Intervention point selection
- Data collection, model access, generation, upload, distribution, reporting, removal 중 어디에 개입하는지 정한다.
- Outcome metric selection
- Classification score가 아니라 actual exposure and remediation outcome을 포함한다.
- Abuse and false-positive analysis
- Defense가 attacker에게 어떤 capability를 주는지, legitimate user에게 어떤 burden을 주는지 본다.
- Domain review
- Technical release 전에 victim-support and policy expert와 risk를 검토한다.
4-3. Engineering notes
1) Publicly available does not mean consented
Web에서 접근 가능한 face image를 identity-conditioned generation dataset으로 사용하는 것은 legal availability와 ethical consent를 혼동할 수 있다. Data provenance record에 consent scope가 포함되어야 한다.
2) Identity matching can be dual-use
Face recognition or similarity search는 victim report triage에 도움이 될 수 있지만, attacker가 target content를 찾고 정리하는 데도 악용될 수 있다. API exposure, rate limit, access logging, result granularity를 함께 설계해야 한다.
3) Detection output should be action-aware
Binary fake label을 public UI에 노출하기 전에 다음을 정해야 한다.
- Remove or reduce distribution인가.
- Human review queue로 보낼 것인가.
- Subject notification이 오히려 harm를 키울 수 있는가.
- Evidence preservation이 필요한가.
- Appeal and correction path가 있는가.
4) Watermark assumptions must be explicit
Watermark defense는 generator가 watermark를 삽입하고 attacker가 이를 완전히 제거하지 않는다는 cooperation assumption이 있다. Open weights or custom pipeline에서는 coverage가 낮아질 수 있다.
5) Research artifact release needs abuse review
Dataset sample, benchmark, trained detector, identity embedding, preprocessing script가 각각 다른 misuse surface를 가진다. Reproducibility와 unrestricted release를 동일시하면 안 된다.
5. Evaluation
5-1. Main results
이 position paper에는 새로운 model benchmark나 detector performance table이 없다. Main empirical evidence는 publication landscape analysis다.
1) AIG-NCII is largely absent from technical implementation
Manual-screened 39편 중 34편은 AIG-NCII를 언급하지 않았고, 5편은 mention에 그쳤으며, AIG-NCII-specific implementation은 0편으로 분류된다.
이 결과가 의미하는 것은 deepfake research가 아무 쓸모가 없다는 것이 아니다. Current benchmark and method가 authenticity-centric problem framing에 집중하고, consent-centric harm를 technical requirement로 operationalize하지 않았다는 뜻이다.
2) Current success metrics can be misaligned
높은 detector AUC는 다음을 보장하지 않는다.
- Harmful content가 빠르게 제거된다.
- Subject의 report burden이 줄어든다.
- Re-upload or re-generation이 어려워진다.
- Authentic NCII victim의 risk가 줄어든다.
- False positive가 consensual creator를 부당하게 제재하지 않는다.
따라서 paper leaderboard가 좋아져도 real-world harm outcome이 정체될 수 있다.
3) Technical intervention can create secondary harm
논문이 강조하는 핵심 case는 다음과 같다.
- Label-only response가 harmful content visibility를 정당화할 수 있다.
- Authenticity confirmation이 victim의 plausible deniability를 약화할 수 있다.
- High-quality forensic tool이 attacker에게 validation tool이 될 수 있다.
- Protective perturbation이 platform transform or adaptive attacker에 쉽게 깨질 수 있다.
- Gated access가 legitimate research and marginalized users에게 disproportionate friction을 줄 수 있다.
좋은 evaluation은 benefit만 아니라 이러한 secondary harm를 함께 측정해야 한다.
5-2. What really matters in the experiments
이 논문의 landscape table을 읽을 때 세 가지를 구분해야 한다.
- Search coverage
- Google Scholar query와 venue/citation filter가 전체 field를 얼마나 대표하는가.
- Mention vs implementation
- Harm를 motivation으로 쓰는 것과 threat model, dataset, metric을 실제로 바꾸는 것은 다르다.
- Absence of work vs impossibility
- Specific implementation이 0편이라는 결과는 technical solution이 불가능하다는 증거가 아니라 research priority gap의 증거다.
AIG-NCII-oriented evaluation을 새로 설계한다면 다음 outcome이 중요하다.
- Time to detection and removal
- Total impressions before intervention
- Re-upload recurrence
- Subject report burden
- False accusation and appeal rate
- Identity matching leakage
- Adaptive attacker success
- Cross-platform portability
- User and survivor assessment of remedy quality
이 논문의 가장 중요한 요구는 dataset category 하나를 추가하는 것이 아니다. Research question의 owner를 viewer에서 subject로 이동시키라는 요구에 가깝다.
6. Limitations
- Position paper이며 proposed intervention을 실험하지 않는다.
- Recommendation의 technical feasibility와 comparative effectiveness는 future work로 남는다.
- Literature review가 exhaustive systematic review는 아니다.
- Google Scholar ranking, venue, citation filter, top-100 screening은 low-visibility or recent work를 놓칠 수 있다.
- Final set 39편은 field 전체를 대표하기에 작을 수 있다.
- Search term and inclusion criteria에 따라 category count가 달라질 수 있다.
- Consent를 content alone에서 infer하기 어렵다.
- Contextual metadata와 report process가 필요하며 privacy and verification burden이 생긴다.
- Prevention techniques는 adaptive attacker에 취약할 수 있다.
- Protective perturbation, watermark, identity shield가 preprocessing or retraining으로 우회될 수 있다.
- Gated access에는 equity and open-science trade-off가 있다.
- High-risk capability를 제한하면서 legitimate research와 beneficial use를 과도하게 막지 않아야 한다.
- Harm-aligned metric이 아직 충분히 operationalized되지 않았다.
- Dignity and autonomy를 단일 scalar로 환원하기 어렵고, impacted people 사이에도 priority가 다를 수 있다.
- Platform and legal context가 지역마다 다르다.
- Reporting, evidence retention, removal, appeal design은 jurisdiction and institution에 의존한다.
- Survivor-centered research에도 representation risk가 있다.
- 소수 참여자의 의견을 전체 피해 경험으로 일반화하거나 repeated consultation burden을 줄 수 있다.
7. My Take
7-1. Why this matters for my work
AI safety and responsible AI에서 metric proxy가 harm target을 대체하는 경우가 많다. Toxicity score, hallucination rate, fake probability가 낮아져도 실제 사용자가 겪는 harm가 줄었다는 보장은 없다.
이 논문은 다음 질문을 technical design review에 넣게 만든다.
- Model이 맞게 분류했을 때 어떤 action이 실행되는가?
- False negative보다 false positive가 더 큰 harm를 만드는 group은 누구인가?
- Dataset을 만들기 위해 다시 harmful content를 수집 and expose하고 있지 않은가?
- Release artifact가 defense보다 abuse capability를 더 쉽게 만들지는 않는가?
- Protected subject가 원하는 outcome은 label, removal, evidence, anonymity 중 무엇인가?
특히 Document AI, retrieval, moderation system에서도 public data와 authorized use를 구분해야 한다. Access 가능성은 consent scope가 아니다.
7-2. Reuse potential
1) Two-axis threat matrix
- Authenticity and consent를 separate label로 둔다.
- Model output에서 알 수 없는 consent는 unknown으로 유지한다.
- Unknown을 safe로 자동 처리하지 않는다.
2) Action-linked evaluation
- Detection score 뒤에 moderation action까지 end-to-end로 평가한다.
- Exposure reduction, review latency, appeal outcome을 포함한다.
3) Identity-risk review
- Face embedding, identity adapter, inpainting, personalization feature를 high-risk capability로 별도 inventory한다.
- Access tier, logging, abuse monitoring을 설계한다.
4) Release minimization
- Full harmful sample 대신 derived statistics or controlled access를 사용한다.
- Reproduction에 꼭 필요하지 않은 identity signal은 제거한다.
5) Domain-partner checkpoint
- Dataset creation 전
- Model release 전
- Platform integration 전
- Incident response design 전
각 단계에서 legal, victim-support, moderation expert review를 둔다.
7-3. Follow-up papers
- The State of Deepfake Detection
- Content Provenance and Authenticity Standards
- Adversarial Perturbations for Protecting Personal Images
- Research on Technology-Facilitated Gender-Based Violence
- Studies of Non-Consensual Intimate Imagery Reporting and Removal
- Responsible Release Practices for Dual-Use AI Systems
8. Summary
- 이 position paper는 deepfake research의 authenticity-centric framing이 AIG-NCII의 consent and dignity harm와 맞지 않을 수 있다고 주장한다.
- Artificiality and consent는 orthogonal하며, synthetic label이 harm를 제거하지 않고 authentic label이 safety를 보장하지 않는다.
- Reviewed 39 papers 중 34편은 AIG-NCII를 언급하지 않았고, 5편은 mention only였으며, specific technical implementation은 0편으로 분류된다.
- Detection, provenance, watermarking은 유용하지만 moderation action, identity protection, access control, prevention과 결합되어야 한다.
- Future research는 model accuracy뿐 아니라 exposure reduction, survivor burden, adaptive misuse, dignity restoration을 evaluation target으로 삼아야 한다.
댓글남기기