Abstract: Sensitive and large text collections pose a practical problem for computational social science: confidential records cannot be sent to commercial language-model services, and per-token pricing makes repeated analysis of very large corpora prohibitive. Small open-weight language models that run on a single workstation address both constraints, keeping text local, fixing model versions for reproducibility,… Continue reading Small Language Models for Sensitive and Large-Scale Text Analysis: Lessons from 1.3 Million Child Welfare Narratives
Author: Editor
Calibrating GenAI for Simulation and Inference in Business Research
Abstract: Large language models offer business researchers two forms of leverage: synthetic decision-makers for social simulations, and predictive labels for large-scale empirical analysis. Yet AI outputs are systematically biased, rarely meeting rigorous statistical or theoretical standards. Calibration, I argue, turns abundant-but-imperfect AI into trustworthy research artifacts, illustrated at two stages of the empirical pipeline. For… Continue reading Calibrating GenAI for Simulation and Inference in Business Research
Vibe Researching with Agentic AI for Social Sciences
Abstract: AI Agents with persistent memory, tool access, and specialist skills can now execute multi-step reasoning across the entire research pipeline, from idea to submission. This is not merely the fourth wave of research automation — following statistical computation, digital trace data, and machine learning — but the first to automate reasoning itself, and it… Continue reading Vibe Researching with Agentic AI for Social Sciences
Using Visual Data in Social Science Research
Abstract: Text analysis has been a hot topic in the social sciences over the past decade. In comparison, the volume of visual data, such as social media images, street-view photos, and digitized maps, has grown even more rapidly, yet methods for analyzing them remain underdeveloped. This talk draws on my recent work to illustrate how… Continue reading Using Visual Data in Social Science Research
Quantifying Narrative Similarity Across Languages
Abstract: How can one understand the spread of ideas across text data? This is a key measurement problem in sociological inquiry, from the study of how interest groups shape media discourse, to the spread of policy across institutions, to the diffusion of organizational structures and institution themselves. To study how ideas and narratives diffuse across… Continue reading Quantifying Narrative Similarity Across Languages
The Silicon Gaze: A Typology of Biases and Inequality in LLMs Through the Lens of Place
Abstract: This paper introduces the concept of the silicon gaze to explain how large language models (LLMs) reproduce and amplify long-standing spatial inequalities. Drawing on a 20.3-million-query audit of ChatGPT, we map systematic biases in the model’s representations of countries, states, cities and neighbourhoods. From these empirics, we argue that bias is not a correctable… Continue reading The Silicon Gaze: A Typology of Biases and Inequality in LLMs Through the Lens of Place
Digital Twins for Social Science Research?
Abstract: “Digital twins” is becoming a paradigm for research and applications in various disciplines. To avoid jumping on the bandwagon without a thorough understanding of what it means and what it takes to address the genuine needs in our investigations, this seminar targets the nature and necessity of developing digital twins for social science research.… Continue reading Digital Twins for Social Science Research?
Factorial Difference-in-Differences
Abstract: We formulate factorial difference-in-differences (FDID) as a research design that extends the canonical difference-in-differences (DID) to settings without clean controls. Such situations often arise when researchers exploit cross-sectional variation in a baseline factor and temporal variation in an event affecting all units. In these applications, the exact estimand is often unspecified and justification for… Continue reading Factorial Difference-in-Differences
Spatial Data Mining to Understand Neighbor Problems and Neighborhood Change
Abstract: Large-scale administrative data collected by municipal governments are increasingly used by researchers to examine urban social phenomena and processes. This seminar will demonstrate how council data from Brisbane, Australia, were utilized to investigate the prevalence of neighbor problems through a GIS-based spatial approach. I will explore how the changing process of our cities such… Continue reading Spatial Data Mining to Understand Neighbor Problems and Neighborhood Change
Luck and Success in Millions of Life Courses
Abstract: This talk probes the role of luck in the determination of success in the life course. We ask whether when people get lucky the trajectory of their life success increasingly diverges from that of their unlucky counterpart. We study this question theoretically using basic models of positive feedback. Empirically we look at the lives… Continue reading Luck and Success in Millions of Life Courses
