跳到正文
Hugging Face:Blog·· 2024-11-04精选AI 评分60

Argilla 2.4 发布:无需代码即可在 Hub 上构建微调与评估数据集

Argilla 2.4: Easily Build Fine-Tuning and Evaluation Datasets on the Hub — No Code Required

AI 导读

Hugging Face 发布 Argilla 2.4,用户可从 Hub 任一公开数据集无代码导入,通过 Argilla UI 定义问题并收集人工反馈,用于构建微调和评估数据集。Hub 现有 230k 数据集可作为起点,部署在 Spaces 上默认启用 Hugging Face OAuth,允许任何 Hub 用户参与标注,也可配置为仅限协作者;后续仍可用 Python SDK 做进一步自定义。

推荐理由

官方介绍了从 Hub 数据集无代码导入并收集人工反馈的完整流程,熟悉数据标注协作场景的读者可据此评估是否迁移现有工作流。

正文 · 原文

Image 1: Hugging Face's logoHugging Face


Back to Articles

Argilla 2.4: Easily Build Fine-Tuning and Evaluation Datasets on the Hub — No Code Required

Published November 4, 2024

Update on GitHub

- [x] Upvote 46

  • Image 2
  • Image 3
  • Image 4
  • Image 5
  • Image 6
  • Image 7
  • +40

Image 8: Natalia Elvira's avatar

Natalia Elvira nataliaElv Follow

Image 9: ben burtenshaw's avatar

ben burtenshaw burtenshaw Follow

Image 10: Daniel Vila's avatar

Daniel Vila dvilasuero Follow

We are incredibly excited to share the most impactful feature since Argilla joined Hugging Face: you can prepare your AI datasets without any code, getting started from any Hub dataset! Using Argilla’s UI, you can easily import a dataset from the Hugging Face Hub, define questions, and start collecting human feedback.

Not familiar with Argilla? Argilla is a free, open-source data-centric tool. Using Argilla, AI developers and domain experts can collaborate and build high-quality datasets. Argilla is part of the Hugging Face family and fully integrated with the Hub. Want to know more? Here’s an intro blog post.

Why is this new feature important to you and the community?

  • The Hugging Face hub contains 230k datasets you can use as a foundation for your AI project.
  • It simplifies collecting human feedback from the Hugging Face community or specialized teams.
  • It democratizes dataset creation for users with extensive knowledge about a specific domain who are unsure about writing code.

Use cases

This new feature democratizes building high-quality datasets on the Hub:

  • If you have published an open dataset and want the community to contribute, import it into a public Argilla Space and share the URL with the world!
  • If you want to start annotating a new dataset from scratch, upload a CSV to the Hub, import it into your Argilla Space, and start labeling!
  • If you want to curate an existing Hub dataset for fine-tuning or evaluating your model, import the dataset into an Argilla Space and start curating!
  • If you want to improve an existing Hub dataset to benefit the community, import it into an Argilla Space and start giving feedback!

How it works

First, you need to deploy Argilla. The recommended way is to deploy on Spaces following this guide. The default deployment comes with Hugging Face OAuth enabled, meaning your Space will be open for annotation contributions from any Hub user. OAuth is perfect for use cases when you want the community to contribute to your dataset. If you want to restrict annotation to you and other collaborators, check this guide for additional configuration options.

Video 1 Once Argilla is running, sign in and click the “Import dataset from Hugging Face” button on the Home page. You can start with one of our example datasets or input the repo id of the dataset you want to use.

In this first version, the Hub dataset must be public. If you are interested in support for private datasets, we’d love to hear from you on GitHub.

Argilla automatically suggests an initial configuration based on the dataset’s features, so you don’t need to start from scratch, but you can add questions or remove unnecessary fields. Fields should include the data you want feedback on, like text, chats, or images. Questions are the feedback you wish to collect, like labels, ratings, rankings, or text. All changes are shown in real time, so you can get a clear idea of the Argilla dataset you’re configuring.

Once you’re happy with the result, click “Create dataset” to import the dataset with your configuration. Now you’re ready to give feedback!

You can try this for yourself by following the quickstart guide. It takes under 5 minutes!

This new workflow streamlines the import of datasets from the Hub, but you can still import datasets using Argilla’s Python SDK if you need further customization.

We’d love to hear your thoughts and first experiences. Let us know on GitHub or the HF Discord!

More Articles from our Blog

Image 11 datasets xet hub ## Streaming datasets: 100x More Efficient * Image 12 * Image 13 * Image 14 * Image 15 * +1 87 October 27, 2025 andito, et. al.

Image 16 data datasets dedupe ## Parquet Content-Defined Chunking * Image 17 kszucs 76 July 25, 2025 kszucs

Community

Edit Preview

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

Comment ·Sign up or log in to comment

- [x] Upvote 46

  • Image 18
  • Image 19
  • Image 20
  • Image 21
  • Image 22
  • Image 23
  • Image 24
  • Image 25
  • Image 26
  • Image 27
  • Image 28
  • Image 29
  • +34

System theme

Company

TOSPrivacyAboutCareers

Website

ModelsDatasetsSpacesPricingDocs

来源:Hugging Face:Blog · huggingface.co