从随机性到信号：一种用于大语言模型可靠测量的贝叶斯潜状态模型 (From Stochasticity to Signal: A Bayesian Latent State Model for Reliable Measurement with LLMs)

Large Language Models (LLMs) are increasingly used to automate classification tasks in business, such as analyzing customer satisfaction from text. However, the inherent stochasticity of LLMs, in terms of their tendency to produce different outputs for the same input, creates a significant measurement error problem that is often neglected with a single round of output, or addressed with ad-hoc methods like majority voting. Such naive approaches fail to quantify uncertainty and can produce biased estimates of population-level metrics. In this paper, we propose a principled solution by reframing LLM variability as a statistical measurement error problem and introducing a Bayesian latent state model to address it. Our model treats the true classification (e.g., customer dissatisfaction) as an unobserved latent variable and the multiple LLM ratings as noisy measurements of this state. This framework allows for the simultaneous estimation of the LLM's false positive and false negative error rates, the underlying base rate of the phenomenon in the population, the posterior probability of the true state for each individual observation, and the causal impact of a business intervention, if any, on the latent state. Through simulation studies, we demonstrate that our model accurately recovers true parameters where naive methods fail. We conclude that this methodology provides a general and reliable framework for converting noisy, probabilistic outputs from LLMs into accurate and actionable insights for scientific and business applications.

翻译：大语言模型（LLMs）在商业领域越来越多地被用于自动化分类任务，例如从文本中分析客户满意度。然而，LLMs固有的随机性——即对相同输入倾向于产生不同输出的特性——造成了显著的测量误差问题，这一问题常因单轮输出而被忽视，或通过多数投票等临时方法处理。此类朴素方法无法量化不确定性，并可能导致对总体水平指标的估计产生偏差。本文通过将LLM变异性重新定义为统计测量误差问题，并引入贝叶斯潜状态模型来解决该问题，提出了一种原则性解决方案。我们的模型将真实分类（例如客户不满）视为未观测的潜变量，并将多次LLM评分视作对该状态的噪声测量。该框架能够同时估计LLM的假阳性与假阴性错误率、现象在总体中的潜在基础发生率、每个个体观测的真实状态后验概率，以及商业干预（若存在）对潜状态的因果影响。通过模拟研究，我们证明该模型能准确还原真实参数，而朴素方法则无法实现。我们总结认为，该方法为将LLM产生的噪声概率输出转化为科学及商业应用中准确且可操作的见解，提供了一个通用且可靠的框架。

相关内容

MoDELS

关注 44

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日