Calibrating GenAI for Simulation and Inference in Business Research

Abstract:

Large language models offer business researchers two forms of leverage: synthetic decision-makers for social simulations, and predictive labels for large-scale empirical analysis. Yet AI outputs are systematically biased, rarely meeting rigorous statistical or theoretical standards. Calibration, I argue, turns abundant-but-imperfect AI into trustworthy research artifacts, illustrated at two stages of the empirical pipeline. For pre-hoc simulation, we translate behavioral theory into structured LLM prompts, calibrate them against limited human data, and validate alignment at outcome and reasoning levels; calibrated agents reproduce the behavioral hypotheses of observational learning and the 11-20 money-request game and generalize out-of-sample. For post-hoc inference, calibrated LLM-augmented double machine learning (Aug-DML) fuses sparse experimental labels with abundant LLM pseudo-labels, reducing causal estimation MAPE by up to 32% over standard DML at 1% labeled data while preserving nominal 95% coverage. Together, they trace an agenda: AI serves business research best not as a replacement for theory or experiments, but as a calibrated complement at well-chosen points in the empirical pipeline.

Speaker:

Prof. Philip Renyu Zhang

Associate Professor

Business School

The Chinese University of Hong Kong

Related Posts