Was 2024 the year of datasets? Is 2025 the year for community-built datasets?
It's exciting to see the progress of many languages in FineWeb-C:
- Total annotations submitted: 41,577
- Languages with annotations: 106
- Total contributors: 363
📖🤗 Machine learning Librarian at Hugging Face
📖🤗 Machine learning Librarian at Hugging Face
Was 2024 the year of datasets? Is 2025 the year for community-built datasets?
It's exciting to see the progress of many languages in FineWeb-C:
- Total annotations submitted: 41,577
- Languages with annotations: 106
- Total contributors: 363
📖🤗 Machine learning Librarian at Hugging Face
Researchers: Want your ML datasets to have more impact? Share them on @huggingface Hub!
✨ Benefits:
• Visibility in the ML community
• Interactive data viewer
• Support for TB-scale datasets
• Integration with @DataPolars @pandas_dev @duckdb and more
https://huggingface.co/blog/researcher-dataset-sharing
📖🤗 Machine learning Librarian at Hugging Face
ColPali is revolutionizing multimodal retrieval. Can we make it even more effective with domain-specific fine-tuning?
Check out my latest blog post, where I create a dataset for fine-tuning a ColPali model for a new domain using an open Vision Language Model.
https://danielvanstrien.xyz/posts/post-with-code/colpali/2024-09-23-generate_colpali_dataset.html
📖🤗 Machine learning Librarian at Hugging Face
Can we search for datasets on the @huggingface Hub based on their content?
> Some datasets lack good documentation 😢
> The dataset viewer preview offers a wealth of information
🤔 How about: query -> dataset based on structure content?
Check out V1: https://huggingface.co/spaces/librarian-bots/huggingface-datasets-semantic-search
📖🤗 Machine learning Librarian at Hugging Face
📖🤗 Machine learning Librarian at Hugging Face
📖🤗 Machine learning Librarian at Hugging Face
Almost ready: search for a @huggingface dataset on the Hub from information in the datasets viewer preview!
Soon, you can find deep-cut datasets even if they don't have a full dataset card (you should still document your datasets!)
📖🤗 Machine learning Librarian at Hugging Face
The @huggingface's Semantic Dataset Search is back in action! Find similar datasets by ID or do a semantic search of dataset cards.
Give it a try:
https://huggingface.co/spaces/librarian-bots/huggingface-datasets-semantic-search
📖🤗 Machine learning Librarian at Hugging Face
📖🤗 Machine learning Librarian at Hugging Face
Is your summer reading list still empty? Curious if an LLM can generate a book blurb you'd enjoy and help build a KTO preference dataset at the same time?
A demo using @huggingface Spaces and @gradio to collect LLM output preferences: https://huggingface.co/spaces/davanstrien/would-you-read-it
📖🤗 Machine learning Librarian at Hugging Face
SPIQA from @Google is a large-scale question-answering dataset centred on figures, tables, and text paragraphs from scientific research papers in various computer science domains.
https://huggingface.co/datasets/google/spiqa
📖🤗 Machine learning Librarian at Hugging Face
HelpSteer2 from @nvidia is an open-source dataset to train top-performing reward models!
- 21,362 samples with annotated attributes
- Attributes: Helpfulness, Correctness, Coherence, Complexity, Verbosity
- Multi-turn prompts
- 88.8% on RewardBench
https://huggingface.co/datasets/nvidia/HelpSteer2
📖🤗 Machine learning Librarian at Hugging Face
Created an "Awesome Synthetic Datasets" list in my ongoing quest to learn more about building synthetic datasets using large language models. Currently includes important tools, datasets, and papers.
Check it out here: https://github.com/davanstrien/awesome-synthetic-datasets
📖🤗 Machine learning Librarian at Hugging Face
Translations from 56 contributors, based on a dataset by 314 community members! These translations will facilitate the creation of evaluations, experimentation with SPIN, building DPO datasets, and more. Interested in contributing to datasets? https://github.com/huggingface/data-is-better-together
📖🤗 Machine learning Librarian at Hugging Face
Who wants to be 100!?
https://github.com/huggingface/data-is-better-together
📖🤗 Machine learning Librarian at Hugging Face
As part of the Multilingual Prompt Evaluation Project (MPEP), we are now automatically exporting the @argilla_io datasets to the @huggingface Hub. We have more than 15 active community-led translation efforts collaborating to enhance datasets for various languages. ❤️
📖🤗 Machine learning Librarian at Hugging Face
📖🤗 Machine learning Librarian at Hugging Face
VISION2UI: A Real-World Dataset with Layout for Code Generation from UI Designs: https://huggingface.co/datasets/xcodemind/vision2ui
📖🤗 Machine learning Librarian at Hugging Face
Doing some rare "front end" work 😬
📖🤗 Machine learning Librarian at Hugging Face
Experimenting with TL;DR summaries for @huggingface datasets using a Chrome plugin.
📖🤗 Machine learning Librarian at Hugging Face
📖🤗 Machine learning Librarian at Hugging Face
Inspired by @Dorialexander@sigmoid.social's French newspaper dataset, I uploaded a smaller dataset of British newspapers to @huggingface Hub. France leads in public domain newspapers, but 2.5B more tokens here!
https://huggingface.co/datasets/biglam/hmd_newspapers
📖🤗 Machine learning Librarian at Hugging Face
📖🤗 Machine learning Librarian at Hugging Face