Hugging Face Collection

Data

I champion open-source ethos. Below are datasets my collaborators and I have built together. Feel free to explore and use them in your work.

Open Research Datasets

Resources for trustworthy AI, content detection, and AI safety.

PADBench

PEFTBackdoorsLLMs

DescriptionA benchmark of benign and backdoored parameter-efficient adapters spanning multiple LLMs, PEFT methods, datasets, and attack strategies.

SourceMultiple LLMs and PEFT configurations
Size13,300 adapters

AIGTBench

AI TextSocial MediaDetection

DescriptionHuman- and AI-generated social media posts covering several widely used language models.

SourceMedium, Quora, Reddit
Size845,497 posts

FragFake

Edited ImagesVLMsLocalization

DescriptionA fine-grained benchmark for detecting and localizing edits from modern image-editing models.

SourceGemini-IG, GoT, MagicBrush, UltraEdit
Size32,326 image-text examples

MGT-Academic

Academic WritingAI TextDetection

DescriptionHuman- and machine-generated academic text spanning STEM, social sciences, and humanities.

SourcearXiv, Wikipedia, Project Gutenberg
Size749K samples ยท 336M+ tokens

Download counts are refreshed from Hugging Face when the page loads.