Hugging Face Recommendations for the U.S. National AI Research Resource

Hugging Face has formally submitted a response to the White House Office of Science and Technology Policy and the National Science Foundation regarding the implementation roadmap for the National Artificial Intelligence Research Resource (NAIRR). The company advocates for a framework that democratizes machine learning by prioritizing ethical oversight, standardized documentation, and accessibility for non-technical researchers.

Prioritizing Technical and Ethical Expertise

NAIRR should appoint advisors who possess both technical expertise and a proven track record of ethical innovation. This dual expertise is necessary to determine what is technically feasible and implementable while ensuring that AI systems do not exacerbate harmful biases or enable malicious use. Hugging Face cites Dr. Margaret Mitchell, Chief Ethics Scientist at Hugging Face, as an example of the type of external advisor needed for this calibration.

Standardizing Model and Data Documentation

To improve accessibility and readability across diverse audiences, NAIRR should establish and provide standardized templates for dataset and system documentation. Hugging Face suggests that Model Cards—a widely adopted documentation structure—serve as a strong template for AI model documentation to ensure consistency and transparency.

Expanding Accessibility for Non-Technical Experts

AI research should be accessible to interdisciplinary and non-technical experts through the provision of educational resources, intuitive interfaces, and low-code or no-code tools. This allows experts from various fields to perform complex tasks, such as training AI models, without requiring deep technical skills. Hugging Face points to its AutoTrain tool as an example of technology that enables users to train, evaluate, and deploy natural language processing (NLP) models regardless of their technical background.

Monitoring and Mitigating Malicious Use

NAIRR must define and continually update a definition of "harm," which should include hate speech, political disinformation, and egregious harmful biases. To mitigate these risks, Hugging Face recommends that NAIRR invest in legal expertise to develop Responsible AI Licenses (RAIL), which provide a mechanism to take action against actors who misuse AI resources.

Empowering Diverse Global Perspectives

To drive responsible innovation, AI tooling and resources must be accessible to researchers across different disciplines and languages. Hugging Face recommends that resources be provided in multiple languages, specifically those most spoken in the U.S. The BigScience Research Workshop—a community of over 1,000 researchers from 60+ countries hosted by Hugging Face and the French government—is cited as a successful model for incorporating diverse global perspectives to build powerful open-source multilingual language models.

Sources