Fetch AI Infrastructure Migration: Consolidating Tools with Hugging Face and AWS
Executive Summary
Fetch, a consumer rewards company serving 18 million active monthly users, has transitioned from relying on third-party AI providers to an in-house AI-powered platform. By leveraging Hugging Face and Amazon Web Services (AWS), Fetch consolidated approximately 15 different AI tools to process over 11 million receipts daily, resulting in a 30% reduction in development time and a 50% reduction in processing latency.
Transition from Third-Party AI to In-House ML
Fetch previously utilized a third-party AI solution to scan and process receipts, which the company described as a "black box." This dependency created business risks and limited the granularity of data insights, preventing Fetch from providing business partners with detailed information on how customers engaged with promotions.
To resolve this, Fetch built its own machine learning (ML) and AI expertise in-house. The company completed the transition in eight months—four months ahead of the original 12-month schedule—by utilizing AWS resources for model training.
Technical Implementation and Tooling
Fetch implemented its new AI infrastructure by integrating several key technologies and programs:
Hugging Face Integration
Fetch engaged with the Hugging Face Expert Acceleration Program via the AWS Marketplace. Hugging Face provided advisory services and knowledge transfer, enabling Fetch engineers to effectively use open-source transformer models to improve entity resolution and semantic search in document AI models.
AWS Infrastructure
- Amazon SageMaker: Used to build, train, and deploy ML models using fully managed infrastructure and workflows.
- AWS Inferentia: Deployed as accelerators to achieve high performance and lower costs for deep learning (DL) inference applications.
Deployment Strategy
To ensure stability during the migration, Fetch utilized a "shadow pipeline." This system reprocessed all 11 million daily receipts in the new ML pipeline for auditing and analytics purposes while the third-party solution remained active, ensuring the new system's accuracy before it became the primary data source.
Key Outcomes and Performance Gains
Moving to a self-managed AI stack has provided Fetch with significant operational and technical improvements:
- Latency Reduction: Processing latency for receipt uploads was cut by 50%, improving the user experience and customer retention.
- Development Efficiency: Development time was reduced by approximately 30% through the guidance of Hugging Face and the use of its open-source resources.
- Training Speed: Model training time was reduced from days or weeks to just a few hours.
- Data Control: By eliminating the "black box" dependency, Fetch gained full control over its data and the ability to extract the specific granularity of insights required by its business partners.