10 분 소요

0. Introduction

Paper link

LLM assistant는 social advice를 자주 요청받는다.

  • 동료가 정말 도우려 한 것인지 묻는다.
  • 친구가 왜 다시 연락했는지 해석해 달라고 한다.
  • 상대의 말이 호의인지 manipulation인지 판단해 달라고 한다.
  • 짧은 message exchange에서 intention or emotion을 추론해 달라고 한다.

이 setting은 standard Theory of Mind benchmark와 다르다. Assistant는 interaction transcript를 직접 보지 않는다. User가 기억하고 선택하고 해석한 narrative만 받는다. Observable event, user belief, emotional framing이 섞여 있고, third party의 true motive는 real world에서 대개 확인할 수 없다.

평가도 어렵다. Assistant가 user에게 공감한 것인지, evidence를 잘 추론한 것인지, biased framing을 그대로 반복한 것인지 구분하기 어렵다. Real social situation에는 verifiable hidden-state label이 없기 때문이다.

Fuse, Framework for User-mediated Social Evaluation은 이 문제를 simulation으로 바꾼다. Target agent에게 hidden motive를 먼저 부여하고 interaction을 생성한다. Simulated user가 이를 debrief한 뒤 evaluated assistant가 motive를 추론한다. Ground truth는 interaction 생성 전에 정해져 있으므로 verifiable하다.

한 줄 요약: Fuse는 hidden motive가 부여된 target agent, interaction participant, simulated user, evaluated assistant를 분리해 user-mediated social reasoning을 ground-truth-known task로 만들고, bias, detail, conversation length가 assistant inference에 미치는 영향을 측정하는 multi-agent evaluation framework다.

이 논문을 지금 볼 가치가 있는 이유는 다음과 같음.

  • Social reasoning을 complete transcript가 아니라 subjective user retelling condition에서 평가한다.
  • Hidden motive를 simulation에서 선할당해 verifiable label을 만든다.
  • User framing bias, narrative detail, multi-turn clarification을 independent factor로 조절한다.
  • Correct, Incorrect, Not Attempted를 분리해 reckless guessing and over-abstention을 같이 본다.
  • 12 LLMs, 21,600 debrief messages, 24,000 human annotations으로 simulation faithfulness를 검증한다.
  • Social advice assistant가 user interpretation을 independent evidence처럼 재포장하는 risk를 보여준다.

1. Problem Setting

1-1. User-mediated social reasoning

Underlying interaction을 $I$, target motive를 $M$, user debrief를 $U$, assistant prediction을 $\hat{M}$이라고 하자.

Real assistant가 보는 것은 $I$가 아니라 $U$다.

\[I,M \rightarrow U \rightarrow \hat{M}\]

User debrief $U$에는 다음이 섞일 수 있다.

  • 관측된 action
  • 누락된 event
  • Emotional language 표현
  • User interpretation
  • Prior belief 정보
  • Leading hypothesis 정보
  • 잘못된 causal attribution

Assistant가 해야 할 일은 user belief를 맞히는 것이 아니다. Available evidence에서 target motive를 추론하거나, evidence가 부족하면 abstain해야 한다.

1-2. Standard social-reasoning benchmark의 한계

많은 Theory of Mind benchmark는 full story or objective narrator를 제공한다. Model은 relevant fact를 직접 읽고 belief, intention, knowledge를 답한다.

하지만 practical social consultation에는 mediation gap이 있다.

  1. Assistant는 event를 직접 보지 못한다.
  2. User가 무엇을 말할지 선택한다.
  3. User는 이미 hypothesis를 가지고 있을 수 있다.
  4. Assistant response는 user에게 independent second opinion처럼 느껴질 수 있다.

Full transcript benchmark가 높아도 mediated narrative에서 robust하다고 보장할 수 없다.

1-3. Ground truth problem

Real social motive는 다음 이유로 verifiable하지 않다.

  • Person이 self-report를 하지 않을 수 있다.
  • Self-report가 truthful하지 않을 수 있다.
  • Motive가 mixed and changing일 수 있다.
  • Observer agreement가 낮을 수 있다.
  • Ethical reason으로 controlled experiment가 어렵다.

Fuse는 realism을 일부 포기하고 identifiability를 얻는다. Motive를 first-class simulation variable로 지정하고, interaction generation이 그 motive를 manifest하도록 만든다.

1-4. Sycophancy와 다른 점

Ordinary sycophancy benchmark는 user의 factual belief or preference에 동의하는지 본다. Fuse는 더 subtle한 setting을 본다.

  • User가 explicit answer를 말하지 않을 수 있다.
  • Framing and selective detail이 hypothesis를 암시한다.
  • Assistant는 evidence-based inference처럼 보이는 answer를 만든다.

따라서 failure는 단순 agreement가 아니라 subjective framing amplification이다.

2. Core Idea

2-1. Ground truth by construction

Scenario template $S$와 two contrasting motives $M_0,M_1$을 준비한다. Target agent는 motive 하나를 hidden instruction으로 받고 other agents와 interaction한다.

\[I \sim p(I \mid S,M)\]

Simulated user agent는 interaction에서 user-visible experience를 바탕으로 debrief를 만든다.

\[U \sim p(U \mid I,B,D)\]
  • $B$: neutral or opposing-belief bias condition
  • $D$: low, medium, high detail level

Evaluated assistant는 $U$만 보고 $M_0$ or $M_1$을 선택하거나 Not Attempted를 낸다.

2-2. ATOMS motive categories

Fuse는 five social-mental-state categories를 사용한다.

  • Desire category
  • Intention category
  • Belief category
  • Emotion category
  • Knowledge category

이 categories는 같은 behavior라도 hidden state가 다를 수 있는 scenario를 만든다. Example로 target이 information을 주지 않는 behavior가 ignorance, concealment intention, negative emotion 중 무엇인지 구분하도록 설계할 수 있다.

2-3. Scenario construction

Dataset structure는 다음과 같다.

  • Scenario templates는 30개
  • Five categories 각각에 six templates
  • Template마다 two contrasting motives
  • Template-motive pair마다 20 interaction realizations
  • Underlying interactions는 총 1,200개

각 interaction에서 debrief condition을 조합한다.

  • Bias conditions는 두 종류
  • Detail levels는 세 종류
  • Debrief variants는 세 종류

따라서 interaction당 18 user messages, total 21,600 evaluation messages가 만들어진다.

2-4. Bias intervention

Neutral condition은 target motive에 대한 leading belief를 최소화한다. Opposing-belief condition은 true motive와 반대되는 user interpretation을 narrative framing에 넣는다.

중요한 점은 evidence 자체와 belief framing을 분리하려는 것이다. Assistant가 factual detail보다 user conclusion에 끌리는지 측정한다.

2-5. Detail intervention

Low, medium, high detail condition은 user message가 포함하는 evidence amount를 바꾼다.

  • Low: sparse description을 제공
  • Medium: additional action and context를 제공
  • High: richer evidence와 interaction detail을 제공

Human and model의 performance curve를 비교하면 model이 human보다 더 많은 explicit detail을 필요로 하는지 볼 수 있다.

2-6. Multi-turn clarification

Single-message evaluation과 별도로 assistant가 clarification question을 묻는 multi-turn setting을 구성한다. Assistant는 new detail을 받고 motive prediction을 revise할 수 있다.

이 experiment는 more turns가 항상 better reasoning을 만드는지 확인한다. Conversation length가 늘어도 new evidence보다 repeated framing이 많으면 performance가 plateau or decline할 수 있다.

3. Architecture / Method

3-1. Overview

Stage Actor Hidden or visible information Output
Scenario setup Framework Template and motive pair Controlled world state
Interaction Target and other agents Target sees assigned motive Social interaction transcript
Debrief User agent User-visible interaction and bias condition Subjective narrative
Evaluation Assistant under test Debrief only Motive prediction or abstention
Scoring Judge and rule Ground-truth motive Correct, Incorrect, Not Attempted

3-2. Target-agent interaction

Target agent는 hidden motive를 받지만 other agents는 이를 보지 못한다. Interaction은 motive를 직접 말하지 않으면서도 그 motive와 일관된 behavior를 드러내야 한다.

Full transcript를 보고도 motive를 구분할 수 없다면 benchmark item 자체가 실패한 것이다. 그래서 paper는 generated interaction이 assigned motive를 실제로 드러내는지 human validation으로 확인한다.

3-3. User debrief

User agent는 모든 turn을 단순 요약하지 않는다. Controlled detail and bias 조건에 따라 participant-mediated account를 만든다.

이 부분이 핵심 interface다. Assistant는 아래 항목을 구분해야 한다.

  • User가 직접 관찰했다고 보고한 것
  • User가 해석하거나 추론한 것
  • Narrative에서 빠진 것
  • 여전히 plausible한 alternative motive

3-4. Prediction and judge

Assistant response는 세 outcome으로 mapping된다.

  1. Correct
  2. Incorrect
  3. Not Attempted

Not Attempted에는 refusal, 선택 없는 explicit uncertainty, non-answer가 paper의 evaluation contract에 따라 포함된다.

LLM judge가 open-ended response를 이 label로 parsing한다. 따라서 judge reliability and model-family bias가 중요하다.

3-5. Motive Success Rate

Paper는 correctness를 보상하면서 cautious abstention에 partial credit을 주기 위해 Motive Success Rate, MSR을 사용한다.

\[\operatorname{MSR}(d) = \frac{ N_{correct} + d N_{not\ attempted} }{N_{total}}\]

Main setting은 $d=0.75$를 사용한다.

  • Incorrect는 0을 받는다.
  • Correct는 1을 받는다.
  • Not Attempted는 0.75를 받는다.

이 metric에는 normative preference가 들어 있다. Confident but wrong inference를 abstention보다 나쁘게 보며, product setting에 따라 다른 $d$가 필요할 수 있다.

4. Training / Data / Recipe

4-1. This is an evaluation framework

Fuse의 main experiment는 new social-reasoning model을 학습하지 않는다. Controlled interaction and user-message data를 만든 뒤 existing LLMs를 평가한다.

따라서 핵심 recipe는 benchmark generation and validation이다.

  1. Scenario template and motive pair를 작성한다.
  2. Multiple target-agent interactions를 생성한다.
  3. Assigned motive가 behavior에 드러나는지 검증한다.
  4. Bias and detail control 아래 user debrief를 생성한다.
  5. Evaluated assistant에 query한다.
  6. Correct, Incorrect, Not Attempted를 판정한다.
  7. Component rates and MSR을 aggregate한다.

4-2. Human validation

Simulation faithfulness는 24,000 human annotations로 검증한다.

여기서는 두 질문이 중요하다.

  • Human annotator가 full interaction transcript에서 assigned motive를 식별할 수 있는가.
  • Mediated condition의 first user message만으로도 human이 motive를 식별할 수 있는가.

Reported result는 full-transcript motive identification 약 96.9%, human first-message correctness 88%, MSR 89.8을 포함한다.

이는 generated interaction이 대체로 motive를 드러내며, mediated task가 어렵지만 풀 수 있는 문제라는 근거가 된다.

4-3. Model evaluation

Paper는 12 LLMs를 평가한다. Model은 같은 user-mediated task를 받고 motive를 추론하거나 abstain해야 한다.

각 model은 아래 세 rate를 함께 보고해야 한다.

  • Correct rate를 보고
  • Incorrect rate를 보고
  • Not Attempted rate를 보고

MSR 하나만 보면 response strategy가 가려질 수 있다. Answer를 거의 하지 않는 model은 low error and low utility를 동시에 가질 수 있고, aggressive model은 high correct and high incorrect가 함께 나타날 수 있다.

4-4. Engineering notes

1) Observable fact and user interpretation should be separate fields

Production assistant prompt and output schema는 아래 항목을 명시적으로 구분해야 한다.

  • Reported event field
  • User interpretation field
  • Model inference field
  • Uncertainty field
  • Missing evidence field

그렇지 않으면 model-generated text가 user belief를 apparently independent evidence로 세탁할 수 있다.

2) Bias condition should preserve facts

Opposing framing experiment는 factual evidence가 comparable할 때만 해석 가능하다. Omitted fact의 차이와 wording의 차이를 별도로 통제해야 한다.

3) Judge disagreement should be logged

Open-ended response에는 mixed confidence or multiple hypotheses가 함께 들어갈 수 있다. One judge label은 brittle할 수 있으므로 multi-judge agreement or human-audit subset이 유용하다.

4) Abstention credit is product-dependent

Relationship advice, HR complaint, clinical context, harmless fiction은 cost matrix가 다르다. Deployment-specific utility analysis 없이 $d=0.75$를 그대로 복사하면 안 된다.

5) Simulator and evaluator family should be diversified

Scenario generation, user narration, response judging에 same model family를 쓰면 shared style and bias가 생길 수 있다. Cross-family generation and judge swap이 중요한 robustness check다.

6) Social advice output should avoid motive certainty

Benchmark가 forced choice를 요구하더라도 deployed assistant는 one account만으로 third-party motive를 알 수 없다고 밝혀야 한다. Evidence and alternative hypotheses를 분리해야 한다.

5. Evaluation

5-1. Overall model-human gap

Strongest reported model인 Gemini 3.7 Flash는 MSR 83.7을 기록한다. Human majority baseline은 correctness 88%, MSR 89.8을 기록한다.

Evaluated model 중 human first-message performance에 도달한 model은 없다. Several models는 incorrect prediction rate가 20%를 넘고, 다른 model은 frequent abstention으로 error를 줄인다.

이 behavior difference 때문에 Correct, Incorrect, Not Attempted를 separately 봐야 한다.

5-2. User mediation cost

Models는 user debrief보다 objective interaction transcript를 받을 때 더 높은 성능을 낸다. 이 gap은 mediation cost를 분리해 보여준다.

User-mediated condition은 두 문제를 추가한다.

  • Information loss
  • Framing contamination

Direct social reasoning capability가 좋아도 these two errors를 robust하게 처리하지 못할 수 있다.

5-3. Bias sensitivity

Opposing-belief framing은 model performance를 일관되게 낮춘다.

Low-abstention models에서 reported MSR drop은 approximately 6.9 to 12.5 points이고, average 7.7 points다. Human drop은 3.6 points다.

Underlying evidence가 다른 motive를 가리켜도 model이 human보다 user의 pre-existing interpretation에 더 민감하다는 의미다.

5-4. More detail helps models more than humans

Average model MSR는 low-detail 74.2에서 high-detail 80.1로 5.9 points 오른다.

Human performance는 88.3에서 91.1로 2.8 points 오른다.

Models는 correct inference에 도달하려면 더 explicit한 evidence가 필요하다. Weaker implicit social prior, poor uncertainty calibration, stronger lexical-cue dependence가 원인일 수 있다.

5-5. Longer conversation is not monotonic improvement

Multi-turn experiment는 approximately turn 4까지 improve한 뒤 plateau or decline하는 pattern을 보인다.

  • Turn 8에서 initial abstention 중 18%만 abstention으로 남는다.
  • Status가 바뀐 case 중 58%는 correct, 24%는 incorrect가 된다.
  • Turn 2 prediction의 81%는 turn 8까지 unchanged다.

More conversation은 uncertainty를 줄일 수 있지만 model을 premature commitment or repeated framing reinforcement로 밀 수도 있다.

5-6. What really matters in the experiments

1) Human baseline is not perfect ground truth

Motive ground truth는 simulation assignment에서 온다. Human performance는 motive가 inferable한지 측정할 뿐, motive label이 ethically or socially complete한지는 보장하지 않는다.

2) MSR ranking and error rate ranking can differ

$0.75$ abstention credit은 cautious model에 substantial score를 준다. Answer coverage의 value가 다른 product는 다른 operating point를 선호할 수 있다.

3) Bias robustness needs evidence-matched pairs

Opposing narrative가 extra false statement or omission을 포함한다면 pure framing effect와 information effect가 섞일 수 있다. Pair construction을 확인해야 한다.

4) Multi-turn improvement must track question quality

Turns count보다 assistant가 어떤 clarification question을 했는지 중요하다. Evidence-seeking question and leading question을 separate metric으로 봐야 한다.

6. Limitations

  1. Synthetic interaction이다.
    • Real relationship의 motive는 mixed, changing, unknowable할 수 있다.
    • Simulated agent는 human보다 motive를 더 clean하게 표현할 수 있다.
  2. Forced contrast는 social reality를 단순화한다.
    • Two candidate motives는 classification을 쉽게 만들고 valid alternative를 억제할 수 있다.
    • Real advice에는 multiple simultaneous hypotheses가 필요한 경우가 많다.
  3. Ground truth는 assigned label이다.
    • Verifiability는 construction으로 얻지만 ecological validity는 제한된다.
  4. MSR의 abstention weight는 normative choice다.
    • $d=0.75$가 universally correct한 것은 아니다.
    • Harm and utility model이 달라지면 ranking도 바뀐다.
  5. LLM judge and simulator bias가 남을 수 있다.
    • Shared model family, style, safety policy가 generation and scoring에 영향을 줄 수 있다.
  6. Scenario and cultural coverage가 제한적이다.
    • Thirty templates로 full interpersonal and cultural variation을 대표할 수 없다.
  7. Benchmark success가 social diagnosis 권한을 주지는 않는다.
    • Good score는 real third party에 대한 confident judgment를 정당화하지 않는다.
    • High-stakes domain에는 uncertainty, consent, provenance, escalation policy가 필요하다.

7. My Take

7-1. Why this matters for my work

Fuse의 가장 중요한 contribution은 social intelligence score가 아니다. Epistemic provenance를 measurable한 evaluation target으로 만든 점이다.

Assistant는 user report를 읽고 answer를 만들지만, new independent evidence를 얻은 것은 아니다. 그런데 fluent response는 user interpretation을 external validation처럼 보이게 할 수 있다.

Reliable assistant는 세 layer를 보존해야 한다.

  1. Observed or reported fact
  2. User interpretation
  3. Model inference

이 boundary가 무너지면 assistant는 evidence를 추가하지 않고 certainty만 추가한다.

7-2. Reuse potential

1) Customer support and complaint analysis

One party’s narrative에서 counterpart intention을 단정하지 않도록 bias and evidence separation test를 만들 수 있다.

2) HR and workplace assistant

Manager, employee, peer report가 서로 다른 framing을 가질 때, assistant가 one-sided motive attribution을 얼마나 amplifies하는지 평가할 수 있다.

3) Agent user simulator evaluation

User simulator가 hidden goal and belief를 갖도록 만들고, assistant가 surface wording보다 latent goal을 recover하는지 볼 수 있다.

4) Calibrated social response schema

Forced motive choice 대신 다음 output을 평가할 수 있다.

  • Supported evidence를 분리
  • Alternative hypotheses를 분리
  • Missing information을 표시
  • Confidence를 표시
  • Recommended clarification question을 제시
  • Harm-sensitive abstention을 허용

7-3. Follow-up papers

  • FANToM: interactive Theory of Mind benchmark
  • Towards Understanding Sycophancy in Language Models 논문
  • Social bias and perspective-taking benchmark 연구
  • LLM assistant의 calibration과 selective prediction 연구
  • Multi-turn information gathering과 clarification-question evaluation 연구

8. Summary

  • Fuse는 user-mediated social reasoning을 hidden-motive simulation으로 평가한다.
  • Target motive를 interaction 전에 지정해 ground truth를 만든다.
  • 30 templates, 1,200 interactions, 21,600 debrief messages로 bias, detail, multi-turn effect를 측정한다.
  • Strongest model MSR 83.7은 human 89.8보다 낮다.
  • Opposing user framing은 model에 human보다 큰 performance drop을 만든다.
  • More detail은 도움이 되지만 longer conversation은 turn 4 이후 monotonic improvement를 보장하지 않는다.
  • 핵심 실무 교훈은 reported fact, user interpretation, model inference를 분리하는 것이다.

댓글남기기