Small Language Models for Sensitive and Large-Scale Text Analysis: Lessons from 1.3 Million Child Welfare Narratives

Abstract: Sensitive and large text collections pose a practical problem for computational social science: confidential records cannot be sent to commercial language-model services, and per-token pricing makes repeated analysis of very large corpora prohibitive. Small open-weight language models that run on a single workstation address both constraints, keeping text local, fixing model versions for reproducibility,… Continue reading Small Language Models for Sensitive and Large-Scale Text Analysis: Lessons from 1.3 Million Child Welfare Narratives

Calibrating GenAI for Simulation and Inference in Business Research

Abstract: Large language models offer business researchers two forms of leverage: synthetic decision-makers for social simulations, and predictive labels for large-scale empirical analysis. Yet AI outputs are systematically biased, rarely meeting rigorous statistical or theoretical standards. Calibration, I argue, turns abundant-but-imperfect AI into trustworthy research artifacts, illustrated at two stages of the empirical pipeline. For… Continue reading Calibrating GenAI for Simulation and Inference in Business Research

Vibe Researching with Agentic AI for Social Sciences

Abstract: AI Agents with persistent memory, tool access, and specialist skills can now execute multi-step reasoning across the entire research pipeline, from idea to submission. This is not merely the fourth wave of research automation — following statistical computation, digital trace data, and machine learning — but the first to automate reasoning itself, and it… Continue reading Vibe Researching with Agentic AI for Social Sciences

Using Visual Data in Social Science Research

Abstract: Text analysis has been a hot topic in the social sciences over the past decade. In comparison, the volume of visual data, such as social media images, street-view photos, and digitized maps, has grown even more rapidly, yet methods for analyzing them remain underdeveloped. This talk draws on my recent work to illustrate how… Continue reading Using Visual Data in Social Science Research

Quantifying Narrative Similarity Across Languages

Abstract: How can one understand the spread of ideas across text data? This is a key measurement problem in sociological inquiry, from the study of how interest groups shape media discourse, to the spread of policy across institutions, to the diffusion of organizational structures and institution themselves. To study how ideas and narratives diffuse across… Continue reading Quantifying Narrative Similarity Across Languages

The Silicon Gaze: A Typology of Biases and Inequality in LLMs Through the Lens of Place

Abstract: This paper introduces the concept of the silicon gaze to explain how large language models (LLMs) reproduce and amplify long-standing spatial inequalities. Drawing on a 20.3-million-query audit of ChatGPT, we map systematic biases in the model’s representations of countries, states, cities and neighbourhoods. From these empirics, we argue that bias is not a correctable… Continue reading The Silicon Gaze: A Typology of Biases and Inequality in LLMs Through the Lens of Place

Digital Twins for Social Science Research?

Abstract: “Digital twins” is becoming a paradigm for research and applications in various disciplines. To avoid jumping on the bandwagon without a thorough understanding of what it means and what it takes to address the genuine needs in our investigations, this seminar targets the nature and necessity of developing digital twins for social science research.… Continue reading Digital Twins for Social Science Research?

Factorial Difference-in-Differences

Abstract: We formulate factorial difference-in-differences (FDID) as a research design that extends the canonical difference-in-differences (DID) to settings without clean controls. Such situations often arise when researchers exploit cross-sectional variation in a baseline factor and temporal variation in an event affecting all units. In these applications, the exact estimand is often unspecified and justification for… Continue reading Factorial Difference-in-Differences

Spatial Data Mining to Understand Neighbor Problems and Neighborhood Change

Abstract: Large-scale administrative data collected by municipal governments are increasingly used by researchers to examine urban social phenomena and processes. This seminar will demonstrate how council data from Brisbane, Australia, were utilized to investigate the prevalence of neighbor problems through a GIS-based spatial approach. I will explore how the changing process of our cities such… Continue reading Spatial Data Mining to Understand Neighbor Problems and Neighborhood Change

Luck and Success in Millions of Life Courses

Abstract: This talk probes the role of luck in the determination of success in the life course. We ask whether when people get lucky the trajectory of their life success increasingly diverges from that of their unlucky counterpart. We study this question theoretically using basic models of positive feedback. Empirically we look at the lives… Continue reading Luck and Success in Millions of Life Courses