콘텐츠로 이동

arXiv 2503.15850 — LLM Uncertainty Quantification Survey

제목: A Survey on Uncertainty Quantification and Confidence Calibration for Large Language Models
저자: Xiaoou Liu, Tiejin Chen, Longchao Da, Chacha Chen, Zhen Lin, Hua Wei
베뉴: KDD 2025 (Proceedings of 31st ACM SIGKDD, August 3–7, 2025, Toronto)
버전: v1 2025-03-20, v2 2025-06-03


핵심 주장

  1. LLM은 classical aleatoric/epistemic 구분으로 다루기 어려운 고유한 불확실성 소스를 가진다
  2. 전통적인 UQ 방법들(MC Dropout, Deep Ensembles, Bayesian)은 LLM 규모에서 계산 비용 문제
  3. 새로운 분류 체계 제안: 2-axis framework

분류 체계 (taxonomy)

Axis 1: Computational efficiency
  - Single-pass (perplexity, token logprob, softmax entropy)
  - Sampling-based (consistency, semantic entropy)

Axis 2: Uncertainty dimensions
  - Input uncertainty     → aleatoric에 대응
  - Reasoning uncertainty → mixed
  - Parameter uncertainty → epistemic에 대응
  - Prediction uncertainty→ mixed

⚠️ 주의: 흔한 오독

"4-dimensional taxonomy"라고 인용되는 경우가 많지만, 실제로는 2-axis framework. 4개 uncertainty dimension은 서로 독립적이지 않고, 논문 본문에서 각각 aleatoric/epistemic/mixed로 다시 분류된다. 즉, 개념적 혁신이 아니라 LLM 문맥의 엔지니어링 정리에 가깝다.


Adversarial Fact-Check 결과

주장 판정 근거
"전통 UQ 방법이 LLM에 부적합" 과장 논문 내에서 conformal prediction을 "distribution-free"라고 긍정 평가
LLM에 novel uncertainty source 존재 부분 지지 현상 자체는 존재하나, classical framework로 환원 가능 (논문 내 인정)
Input ambiguity/reasoning divergence/decoding stochasticity = beyond classical 논쟁적 대립 논문 2311.08718: clarification ensemble으로 classical로 환원 가능 주장

실무 함의

  • Production에서 실제 쓰이는 것: temperature scaling (GPT-4, Gemini, DeepSeek 기술 보고서 명시), token entropy, conformal prediction
  • 논문의 "inadequacy" 수사는 novel taxonomy 기여를 정당화하기 위한 수사적 포지셔닝에 가깝다
  • 비판적 독자는 "traditional methods struggle at scale"을 "traditional methods fail"로 오독하지 말 것

참고 인용

  • Hüllermeier & Waegeman (2021) — aleatoric/epistemic 분류 원본
  • 대립 taxonomy: arXiv 2603.24967, 2505.07309, 2505.22655 (consensus 없음)