arXiv 2503.15850 — LLM Uncertainty Quantification Survey¶
제목: A Survey on Uncertainty Quantification and Confidence Calibration for Large Language Models
저자: Xiaoou Liu, Tiejin Chen, Longchao Da, Chacha Chen, Zhen Lin, Hua Wei
베뉴: KDD 2025 (Proceedings of 31st ACM SIGKDD, August 3–7, 2025, Toronto)
버전: v1 2025-03-20, v2 2025-06-03
핵심 주장¶
- LLM은 classical aleatoric/epistemic 구분으로 다루기 어려운 고유한 불확실성 소스를 가진다
- 전통적인 UQ 방법들(MC Dropout, Deep Ensembles, Bayesian)은 LLM 규모에서 계산 비용 문제
- 새로운 분류 체계 제안: 2-axis framework
분류 체계 (taxonomy)¶
Axis 1: Computational efficiency
- Single-pass (perplexity, token logprob, softmax entropy)
- Sampling-based (consistency, semantic entropy)
Axis 2: Uncertainty dimensions
- Input uncertainty → aleatoric에 대응
- Reasoning uncertainty → mixed
- Parameter uncertainty → epistemic에 대응
- Prediction uncertainty→ mixed
⚠️ 주의: 흔한 오독¶
"4-dimensional taxonomy"라고 인용되는 경우가 많지만, 실제로는 2-axis framework. 4개 uncertainty dimension은 서로 독립적이지 않고, 논문 본문에서 각각 aleatoric/epistemic/mixed로 다시 분류된다. 즉, 개념적 혁신이 아니라 LLM 문맥의 엔지니어링 정리에 가깝다.
Adversarial Fact-Check 결과¶
| 주장 | 판정 | 근거 |
|---|---|---|
| "전통 UQ 방법이 LLM에 부적합" | 과장 | 논문 내에서 conformal prediction을 "distribution-free"라고 긍정 평가 |
| LLM에 novel uncertainty source 존재 | 부분 지지 | 현상 자체는 존재하나, classical framework로 환원 가능 (논문 내 인정) |
| Input ambiguity/reasoning divergence/decoding stochasticity = beyond classical | 논쟁적 | 대립 논문 2311.08718: clarification ensemble으로 classical로 환원 가능 주장 |
실무 함의¶
- Production에서 실제 쓰이는 것: temperature scaling (GPT-4, Gemini, DeepSeek 기술 보고서 명시), token entropy, conformal prediction
- 논문의 "inadequacy" 수사는 novel taxonomy 기여를 정당화하기 위한 수사적 포지셔닝에 가깝다
- 비판적 독자는 "traditional methods struggle at scale"을 "traditional methods fail"로 오독하지 말 것
참고 인용¶
- Hüllermeier & Waegeman (2021) — aleatoric/epistemic 분류 원본
- 대립 taxonomy: arXiv 2603.24967, 2505.07309, 2505.22655 (consensus 없음)