~/AI DATASETS/china-releases-chinese-internet-basic-corpus-4-0-with-120gb-of-data

China Releases Chinese Internet Basic Corpus 4.0 with 120GB of Data

The Cyber Security Association of China has officially released the Chinese Internet Basic Corpus 4.0, offering 120GB of trusted data for artificial intelligence and large language model training. Created under the guidance of the Cyberspace Administration of China, the dataset was developed in collaboration with major organizations including Baidu, iFlytek, and Zhihu. High-quality, curated Chinese text datasets are crucial for closing the data gap between Chinese and English LLM training ecosystems. By expanding the supply of verified open data, this initiative helps establish a standardized data foundation for domestic AI development and safety alignment. The 120GB corpus is accessible via the Chinese Internet Corpus Resource Platform following user registration and authentication procedures. The dataset underwent strict deduplication and processing measures to integrate data contributed by participating tech enterprises and research institutes.

## BACKGROUND

Training performant large language models requires access to vast quantities of sanitized text to prevent issues like toxic content or model hallucination. To address data scarcity in Chinese NLP, Chinese cybersecurity agencies and industry leaders established a co-construction and sharing mechanism for internet corpora. Previous releases (versions 1.0 through 3.0) incrementally built out this public resource to supply reliable pre-training material for AI researchers.

## REFERENCES

## KEYWORDS

#AI Datasets#LLM Training#Natural Language Processing#Chinese NLP

$ subscribe --daily

China Releases Chinese Internet Basic Corpus 4.0 with 120GB of Data | Daily News