|
Best Practices for Large-Scale Image Datasets? (between WebDataset and Parquet)
|
|
4
|
957
|
September 6, 2026
|
|
Open-source tool for detecting issues in robot-learning datasets
|
|
5
|
108
|
September 6, 2026
|
|
Can someone with a Chinese Baidu NetDisk account help me download this dataset?
|
|
5
|
820
|
September 5, 2026
|
|
[Commercial Dataset] 11,000 Hours of Production-Ready Egyptian & Emirati Call Center Audio (Dual-Channel, 48kHz/16-bit)
|
|
0
|
9
|
September 4, 2026
|
|
Open call: test your agent memory layer on an adversarial coding benchmark
|
|
2
|
85
|
September 4, 2026
|
|
Save a dataset while it's being created by from_generator()
|
|
2
|
80
|
September 4, 2026
|
|
[Quota Increase Request] Academic Dataset Storage for [TUTTI]
|
|
3
|
95
|
September 3, 2026
|
|
Reliability check on my own dataset's annotation layer: five machine raters, one definition, answers from 0 to 78
|
|
1
|
63
|
September 2, 2026
|
|
Datasets >= 5 ruins sharding when shuffle() is called
|
|
1
|
73
|
September 1, 2026
|
|
JPEG XL images in the dataset
|
|
1
|
67
|
August 30, 2026
|
|
Dataset Intelligence for Robotics
|
|
1
|
66
|
August 28, 2026
|
|
What are the best practices for detecting and fetching deltas from a dataset?
|
|
4
|
115
|
August 28, 2026
|
|
What are the biggest challenges you face when sourcing high-quality, domain-specific datasets for training and evaluating AI models?
|
|
1
|
55
|
August 28, 2026
|
|
Complete personal archive of Tsiolkovsky (51k sheets) — with recognition accuracy actually measured
|
|
0
|
42
|
August 25, 2026
|
|
New Dataset Preview - Celtic Constellation Reference Sessions — aligned ensemble stems for Celtic instruments
|
|
0
|
46
|
August 25, 2026
|
|
Open 78-card tarot meaning dataset for symbolic NLP and multilingual experiments
|
|
1
|
87
|
August 23, 2026
|
|
A free, monthly-refreshed compliance matrix of all 4,893 Indic-language datasets on the Hub, showing that 65% declare no license — so teams can avoid licensing traps (and missing-tag repos) before they train on them
|
|
2
|
63
|
August 23, 2026
|
|
Script-Generated Visual Datasets with Reproducible Metadata and Adobe Generator Tools
|
|
1
|
66
|
August 20, 2026
|
|
Why is there almost no manipulation data for agriculture?
|
|
2
|
132
|
August 20, 2026
|
|
UncertaintyGym: A Benchmark for LLM Epistemic Calibration and Uncertainty Expression
|
|
4
|
114
|
August 18, 2026
|
|
Human Vocality Primitives — full preview set now available (17 categories)
|
|
0
|
49
|
August 18, 2026
|
|
Feature request httpx2
|
|
1
|
87
|
August 15, 2026
|
|
📢 New Demo Dataset: “DeFi-Behavior 0728” – real on-chain trading episodes, already labeled
|
|
1
|
135
|
August 13, 2026
|
|
[SEEKING] Indic Document Dataset (India) — Invoices, Receipts, Utility Bills, Payment Advices, Packing Lists, Commercial Invoices, Credit Notes
|
|
5
|
196
|
August 7, 2026
|
|
Is selling datasets way harder than building them? Or is it just me?
|
|
2
|
158
|
August 7, 2026
|
|
Same model, up to 4.66x different price — full Inference Providers pricing matrix
|
|
0
|
54
|
August 1, 2026
|
|
cerebras/SlimPajama-627B unexpectedly returns 401/404
|
|
1
|
146
|
August 1, 2026
|
|
Organization Admin needs to urgently delete two datasets, Settings option missing
|
|
1
|
95
|
July 30, 2026
|
|
Fastdedup: Rust-based dataset deduplication — benchmarks on FineWeb sample-10BT
|
|
3
|
185
|
July 28, 2026
|
|
Training LLM model for asking questions
|
|
5
|
473
|
July 28, 2026
|