YODAS v3, a 1.1 million-hour speech dataset, is available on Hugging Face
William Chen says YODAS v3, a speech training dataset, is now available on Hugging Face under the CC-BY-3.0 license. It contains 1.1 million hours of audio and covers more than 100 languages.
Chen describes it as the biggest audio dataset ever, and the first at this scale to include stereo audio at 48 kHz along with timestamped transcripts and translations. Details are on the ESPnet blog post about YODAS v3.