The production of large-scale political datasets typically demands extracting structured facts from vast piles of unstructured documents or web sources, a task that traditionally relies on expensive human experts and remains prohibitively difficult to automate at scale. We leverage Large Language Models (LLMs) to automate the extraction of multi-dimensional elite biographies. We propose a two-stage Synthesis-Coding framework: an upstream synthesis stage that uses recursive agentic LLMs to search, filter, and curate biographies from heterogeneous web sources, followed by a downstream coding stage that maps curated biography into structured dataframes.
@unpublished{zhu2025agentic,title={Agentic Framework for Political Biography Extraction},author={Zhu, Yifei and Yang, Songpo and Zhu, Jiangnan and Jiang, Junyan},year={2026},note={Invited to resubmit, American Journal of Political Science},doi={10.48550/arXiv.2603.18010},}
Publications
PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts
Yifei Zhu
In The 64th Annual Meeting of the Association for Computational Linguistics, 2026
Large Reasoning Models (LRMs) embedded in agentic frameworks have transformed information retrieval from static, long-context question answering into open-ended exploration. Yet real-world use requires models to discover and synthesize "long-tail" facts from dispersed sources, a capability that remains under-evaluated. We introduce PolitNuggets, a multilingual benchmark for agentic information synthesis via constructing political biographies for 400 global elites, covering over 10000 political facts. We standardize evaluation with an optimized multi-agent system and propose FactNet, an evidence-conditional protocol that scores discovery, fine-grained accuracy, and efficiency.
@inproceedings{zhu2025politnuggets,title={PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts},author={Zhu, Yifei},booktitle={The 64th Annual Meeting of the Association for Computational Linguistics},year={2026},note={San Diego, California, United States},doi={10.48550/arXiv.2605.14002},}
Using Silicon Predictions to Augment Rather Than Replace Surveys
@unpublished{wei2025silicon,title={Using Silicon Predictions to Augment Rather Than Replace Surveys},author={Wei, Lai and Zhu, Yifei},year={2026},note={Conditionally accepted, Sociological Science},}
Mobilizing the State Without Crushing Local Knowledge: Evidence From China’s Poverty-Alleviation Campaign
@article{zang2026mobilizing,title={Mobilizing the State Without Crushing Local Knowledge: Evidence From China's Poverty-Alleviation Campaign},author={Zang, Leizhen and Zhu, Yifei and Roth, Antoine},journal={Public Administration and Development},year={2026},pages={1--14},doi={10.1002/pad.70112},}
Book
Chinese-Language Book on Large Language Models for Social Science
@book{zhu2026llmbook,title={Chinese-Language Book on Large Language Models for Social Science},author={Zhu, Yifei and others},publisher={Peking University Press},year={2026},note={Forthcoming; coauthored},}
Working Papers
Power Incubator: How An Authoritarian Anticorruption Agency Facilitates Power Sharing?
@unpublished{zhu2025power,title={Power Incubator: How An Authoritarian Anticorruption Agency Facilitates Power Sharing?},author={Zhu, Yifei and Zhu, Jiangnan},year={2026},note={Working paper},}
Party Organizer: The Selection Pattern of POD Heads from 1949-2025
Junqi Feng, Sidi Huang, Jiangnan Zhu, and Yifei Zhu
@unpublished{feng2025party,title={Party Organizer: The Selection Pattern of POD Heads from 1949-2025},author={Feng, Junqi and Huang, Sidi and Zhu, Jiangnan and Zhu, Yifei},year={2026},note={Working paper},}
A Consensus-Aware Framework with Persona-Seeded LLM Coders: An Application on Career Mobility Networks
Liudan Su, Aoran Cheng, Miner Ye, Jiayi Zeng, and Yifei Zhu
Traditional sociological thinking presumes hidden social structure resides within social norms, values, and common sense. We operationalize this idea as a computational social science measurement problem: can we recover population-level perceptions of social hierarchy from text descriptions in a way that is empirically validated by observed behavior? We cast this as an expert coding task: given an occupation description, a coder assigns a perceived-prestige score that captures shared social meaning and status ordering. Building on a consensus-awareness view of shared perceptions, we construct a framework with persona-seeded LLM coders whose generated perceived occupation prestige scores can predict occupational mobility structure.
@unpublished{su2025consensus,title={A Consensus-Aware Framework with Persona-Seeded LLM Coders: An Application on Career Mobility Networks},author={Su, Liudan and Cheng, Aoran and Ye, Miner and Zeng, Jiayi and Zhu, Yifei},year={2026},note={Working paper},}
January 2026 - New Horizons for AI Research in the Social Sciences and Humanities Conference, University of Chicago Hong Kong Campus
Conference Talks
July 2026 - The 64th Annual Meeting of the Association for Computational Linguistics, San Diego — PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts
September 2025 - APSA Annual Meeting, Vancouver — Agentic Framework for Political Biography Extraction
May 2025 - Fudan-UCSB Young Talented Scholars Conference, San Diego — Agentic Framework for Political Biography Extraction
September 2024 - APSA Annual Meeting, Philadelphia — Power Incubator: How An Authoritarian Anticorruption Agency Facilitates Power Sharing?
April 2024 - Hong Kong Quantitative Social Science Workshop, Hong Kong — Power Incubator: How An Authoritarian Anticorruption Agency Facilitates Power Sharing?
November 2022 - Lien Development Conference, Singapore