Hugging Face Multimodal Project Ethical Charter
Hugging Face has introduced an ethical charter for its multimodal learning project to ensure that ethical considerations are integrated into the research lifecycle from the outset. This framework aims to mitigate risks such as algorithmic bias, data privacy violations, and malicious use by formalizing the values and content policies that govern the project's development.
Purpose and Framework of the Ethical Charter
The multimodal learning group at Hugging Face developed this charter to move ethical considerations from an afterthought to a core component of the machine learning lifecycle. By defining these principles at the project's inception, the team intends to ground all technical choices in a clear set of values and goals.
Key objectives of the charter include:
- Early Feedback: By being transparent about decisions and team roles, Hugging Face seeks community feedback early enough to implement meaningful changes.
- Collaborative Development: The document was created by Hugging Face researchers and engineers with input from experts in personal privacy, data governance, and ethics operationalization.
- Iterative Evolution: The charter is a work in progress as of May 2022, with updates tracked via GitHub to reflect the evolving nature of "ethical AI."
Content Policy and Prohibited Use Cases
Hugging Face has identified specific misuses of multimodal technologies that the project aims to prevent. The content policy prohibits the following:
- Harmful Content: The promotion of violence, harassment, bullying, hate, and discrimination against identity subpopulations based on race, gender, age, ability, LGBTQA+ orientation, religion, education, or socioeconomic status.
- Legal and Rights Violations: Any violation of copyrights, privacy, human rights, cultural rights, fundamental rights, or binding laws.
- PII Generation: The generation of personally identifiable information.
- Misinformation: The generation of false information intended to harm or trigger others, or produced without accountability.
- High-Risk Domain Application: The incautious use of models in sectors that could fundamentally damage lives, specifically medical, legal, finance, and immigration domains.
Core Ethical Values
The project is guided by five primary values designed to ensure accountability and fairness:
- Transparency: Openness regarding intent, data sources, tools, and decisions to allow the community to identify weak points and hold the team accountable.
- Openness and Reproducibility: Sharing precise descriptions of data, tools, and experimental conditions. Research artifacts and model checkpoints must be accessible to all without discrimination.
- Fairness: Defined as the equal treatment of all human beings, requiring the monitoring and mitigation of biases related to race, gender, disabilities, and sexual orientation. This includes conducting bias reviews on both training data and model outputs.
- Self-Criticism: A commitment to avoiding hype and overclaiming, while constantly seeking better strategies for curating and filtering training data.
- Credit Attribution: Respecting the work of others through proper licensing and credit.
Managing Value Conflicts
Hugging Face acknowledges that ethical values can conflict in practice. For example, the goal of sharing open and reproducible work may conflict with the need to respect individual privacy or ensure fairness. The team emphasizes that these tensions require a case-by-case analysis of risks and benefits to determine the best course of action.