← Back to home

Data & Code

Reliable Large Language Models

  • PARADEDataset Paraphrase identification that requires computer science domain knowledge: about 10,000 labeled pairs across 788 CS entities. PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain Knowledge · EMNLP 2020 Mirrored in the tasksource collection on Hugging Face
  • Awesome LLM Self-ImprovementResource A curated, maintained list of papers on inference-time self-improvement for large language models. Companion to our survey on LLM inference-time self-improvement · arXiv 2412.14352

Trustworthy and Efficient Recommendation

  • Mamba4RecCode Reference implementation of selective state space models for efficient sequential recommendation. Mamba4Rec · RelKD@KDD 2024 Best Paper Award Independently reimplemented by others outside our group
  • User-Generated Item Lists (Goodreads, Spotify, Zhihu)Dataset Three collected corpora of user-curated lists: books, songs, and answers, with de-identified raw data and preprocessed splits. A Hierarchical Self-Attentive Model for Recommending User-Generated Item Lists · CIKM 2019 Reused as benchmarks in later list recommendation work
  • List ContinuationCode + Data Data and code for continuing user-generated item lists. Consistency-Aware Recommendation for User-Generated Item List Continuation · WSDM 2020

AI in the Public Interest

  • DisastIRBenchmark An information retrieval benchmark for disaster management: about 240,000 passages and 9,600 queries across 48 tasks spanning six search intents and eight disaster types. DisastIR: A Comprehensive Information Retrieval Benchmark for Disaster Management · EMNLP Findings 2025
  • DisastQABenchmark A question answering benchmark for disaster management: 3,000 verified questions (2,000 multiple-choice and 1,000 open-ended), with a keypoint-based protocol for evaluating open-ended answers. DisastQA: A Comprehensive Benchmark for Question Answering Evaluation in Disaster Management · ACL Findings 2026
  • DMRetrieverModels + Code A family of dense retrieval models for disaster management, from 33M to 7.6B parameters. Checkpoints are on Hugging Face, along with the training data. DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management · ACL 2026

Social Media Mining

  • Social Honeypot DatasetDataset Over 22,000 content polluters and 19,000 legitimate Twitter users, with more than 5.6 million tweets, collected by social honeypots over seven months. A labeled subset is available in the Botometer Bot Repository (as “caverlee-2011”). Seven Months with the Devils: A Long-Term Study of Content Polluters on Twitter · ICWSM 2011 Used to train Botometer, a widely used social bot detection tool
  • Microblog Location DatasetDataset Twitter users with city-level home locations, plus a test set with device-reported coordinates, for content-based geo-location. Hosted on the Internet Archive. You Are Where You Tweet: A Content-Based Approach to Geo-locating Twitter Users · CIKM 2010 Test of Time Award A standard benchmark for content-based Twitter geo-location
  • Location Sharing Services DatasetDataset Check-ins and user data from location-sharing services, collected September 2010 – January 2011. Exploring Millions of Footprints in Location Sharing Services · ICWSM 2011
  • Geo-Tagged Hashtag DatasetDataset About 99,000 hashtags and 21 million geo-tagged, time-stamped occurrences for studying how memes spread across space and time. Spatio-Temporal Dynamics of Online Memes · WWW 2013

Speech and Spoken Language

  • DRESBenchmark An evaluation framework for benchmarking language models on disfluency removal. Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
  • Z-ScoresCode A span-level, linguistically grounded metric for evaluating disfluency removal. Z-Scores · ICASSP 2026
  • Syn-WSSECode A framework for generating realistic whole-word speech substitution errors at controlled rates. INTERSPEECH 2025
  • Comparing ASR SystemsCode Code and analysis comparing Google ASR and WhisperX on human-annotated podcast episodes and over 82,000 episodes at scale. INTERSPEECH 2024