Data & Code
Reliable Large Language Models
- PARADEDataset Paraphrase identification that requires computer science domain knowledge: about 10,000 labeled pairs across 788 CS entities. PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain Knowledge · EMNLP 2020 Mirrored in the tasksource collection on Hugging Face
- Awesome LLM Self-ImprovementResource A curated, maintained list of papers on inference-time self-improvement for large language models. Companion to our survey on LLM inference-time self-improvement · arXiv 2412.14352
Trustworthy and Efficient Recommendation
- Mamba4RecCode Reference implementation of selective state space models for efficient sequential recommendation. Mamba4Rec · RelKD@KDD 2024 Best Paper Award Independently reimplemented by others outside our group
- User-Generated Item Lists (Goodreads, Spotify, Zhihu)Dataset Three collected corpora of user-curated lists: books, songs, and answers, with de-identified raw data and preprocessed splits. A Hierarchical Self-Attentive Model for Recommending User-Generated Item Lists · CIKM 2019 Reused as benchmarks in later list recommendation work
- List ContinuationCode + Data Data and code for continuing user-generated item lists. Consistency-Aware Recommendation for User-Generated Item List Continuation · WSDM 2020
AI in the Public Interest
- DisastIRBenchmark An information retrieval benchmark for disaster management: about 240,000 passages and 9,600 queries across 48 tasks spanning six search intents and eight disaster types. DisastIR: A Comprehensive Information Retrieval Benchmark for Disaster Management · EMNLP Findings 2025
- DisastQABenchmark A question answering benchmark for disaster management: 3,000 verified questions (2,000 multiple-choice and 1,000 open-ended), with a keypoint-based protocol for evaluating open-ended answers. DisastQA: A Comprehensive Benchmark for Question Answering Evaluation in Disaster Management · ACL Findings 2026
- DMRetrieverModels + Code A family of dense retrieval models for disaster management, from 33M to 7.6B parameters. Checkpoints are on Hugging Face, along with the training data. DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management · ACL 2026
Social Media Mining
- Social Honeypot DatasetDataset Over 22,000 content polluters and 19,000 legitimate Twitter users, with more than 5.6 million tweets, collected by social honeypots over seven months. A labeled subset is available in the Botometer Bot Repository (as “caverlee-2011”). Seven Months with the Devils: A Long-Term Study of Content Polluters on Twitter · ICWSM 2011 Used to train Botometer, a widely used social bot detection tool
- Microblog Location DatasetDataset Twitter users with city-level home locations, plus a test set with device-reported coordinates, for content-based geo-location. Hosted on the Internet Archive. You Are Where You Tweet: A Content-Based Approach to Geo-locating Twitter Users · CIKM 2010 Test of Time Award A standard benchmark for content-based Twitter geo-location
- Location Sharing Services DatasetDataset Check-ins and user data from location-sharing services, collected September 2010 – January 2011. Exploring Millions of Footprints in Location Sharing Services · ICWSM 2011
- Geo-Tagged Hashtag DatasetDataset About 99,000 hashtags and 21 million geo-tagged, time-stamped occurrences for studying how memes spread across space and time. Spatio-Temporal Dynamics of Online Memes · WWW 2013
Speech and Spoken Language
- DRESBenchmark An evaluation framework for benchmarking language models on disfluency removal. Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
- Z-ScoresCode A span-level, linguistically grounded metric for evaluating disfluency removal. Z-Scores · ICASSP 2026
- Syn-WSSECode A framework for generating realistic whole-word speech substitution errors at controlled rates. INTERSPEECH 2025
- Comparing ASR SystemsCode Code and analysis comparing Google ASR and WhisperX on human-annotated podcast episodes and over 82,000 episodes at scale. INTERSPEECH 2024