upgini/upgini

Data search & enrichment library for Machine Learning → Easily find and add relevant features to your ML & AI pipeline from hundreds of public and premium external data sources, including open & commercial LLMs

What it solves

Upgini is a low-code library designed to solve the time-consuming process of finding and integrating external data and features into machine learning pipelines. It aims to replace manual data wrangling and hypothesis testing with an automated system that identifies only the features that actually improve a model's accuracy, rather than just those that are correlated with the target variable.

How it works

Upgini acts as an intelligent data search engine. It connects to hundreds of public, community, and premium external data sources. Using a combination of Large Language Models (LLMs), Graph Neural Networks (GNNs), and Recurrent Neural Networks (RNNs), it automatically generates and optimizes an optimal set of ML features. It can also augment search keys (like postal codes or IP addresses) to broaden the search across available sources and provides a Scikit-learn-compatible interface for seamless integration into existing pipelines.

Who it’s for

It is built for data scientists and ML engineers working with tabular data on supervised learning tasks, including binary/multiclass classification, regression, and time-series prediction.

Highlights

  • Automated Feature Discovery: Finds features that specifically provide accuracy uplift for the target model.
  • Multi-Source Integration: Accesses public, community-shared, and premium data providers across 239 countries.
  • Search Key Augmentation: Automatically fills in missing search keys to maximize data retrieval.
  • Stability Verification: Checks accuracy gains on out-of-time intervals and verification datasets to mitigate dependency risks.
  • Low-Code Interface: Offers a Scikit-learn-compatible Python library and a drag-and-drop search UI.

Related

  • Project
  • Project
  • Project
  • Project
  • Project