2101

AirSim: a high-fidelity visual and physical simulator for autonomous vehicles and AI research

An open-source simulator for drones and cars built on Unreal Engine, designed as a platform for AI research in deep learning, computer vision, and reinforcement learning.

2102

MaaAssistantArknights: an image-recognition based automation assistant for Arknights that automates daily tasks and base management

An automation assistant for Arknights that uses image recognition and deep learning to automate daily tasks, base management, and resource farming.

2103

vit-pytorch: a comprehensive collection of Vision Transformer (ViT) and its variants implemented in PyTorch

A comprehensive PyTorch library implementing the original Vision Transformer (ViT) and dozens of its advanced variants for image classification.

2104

label-studio: a multi-modal open-source data labeling tool with ML-assisted pre-labeling and active learning

An open-source data labeling tool that allows users to annotate audio, text, images, and video to prepare or improve training data for machine learning models.

2105

espnet: a comprehensive end-to-end speech processing toolkit for ASR, TTS, and spoken language understanding

An end-to-end speech processing toolkit that provides a unified framework for ASR, TTS, speech translation, and enhancement using PyTorch.

2106

leon: an open-source personal AI assistant with agentic execution and local-first privacy

An open-source personal AI assistant that uses tools, memory, and agentic execution to perform tasks while supporting local AI models for privacy.

2107

Cursor iOS App Installation Changes Account Privacy Settings

Installing the Cursor iOS app can irreversibly switch a user's account from Privacy Mode (Legacy) to a less restrictive Privacy Mode, removing the access to the previous setting.

2108

PlayStation Store Removes Studio Canal Movies from User Libraries Without Refunds

Sony has removed hundreds of Studio Canal movies from PlayStation Store user accounts due to licensing agreements, sparking a debate over the legal nature of digital ownership.

2109

US Supreme Court Rules Geofence Warrants Require Constitutional Protections

The US Supreme Court has ruled that geofence warrants must adhere to Fourth Amendment protections, rejecting the argument that opting into location services waives a user's privacy rights.

2110

Tidal AI Policy: Demonetization and Labeling of AI-Generated Music

Tidal has introduced a new AI policy that demonetizes 100% AI-generated music and implements mandatory labeling to protect human artists and prevent fraudulent impersonations.

2111

HunyuanVideo-1.5: a lightweight 8.3B parameter video generation model for high-quality synthesis on consumer GPUs

A lightweight 8.3B parameter video generation model that enables high-quality text-to-video and image-to-video synthesis on consumer-grade GPUs.

2112

OpenMontage: an agentic video production system that orchestrates research, scripting, and editing into a full production pipeline

An open-source, agentic video production system that orchestrates research, scripting, asset generation, and editing to create professional videos from plain-language prompts.

2113

Ornith-1.0 Release: Self-Improving Open-Source Models for Agentic Coding

Ornith-1.0 is a family of MIT-licensed, self-improving open-source models (9B to 397B) optimized for agentic coding and tool-calling, post-trained on Gemma 4 and Qwen 3.5.

2114

Outer Shell: A Native Graphical Shell for SSH

Outer Shell introduces a browser-based graphical shell for remote servers that leverages SSH for security and Unix domain sockets for private, native app communication.

2115

Rocket Lab Acquires Iridium

Rocket Lab has acquired Iridium, integrating a profitable satellite communications network to secure a baseline of regular launches and gain critical spectrum assets.

2116

Samsung, SK Hynix, and Micron Sued in US for Memory Price Fixing

Samsung, SK Hynix, and Micron face a US lawsuit alleging they coordinated the reduction of DRAM supply and the discontinuation of older memory standards to inflate prices.

2117

ZLUDA 6 Release: Running Unmodified CUDA Applications on Non-NVIDIA GPUs

ZLUDA 6 introduces support for 32-bit PhysX, basic texture support for Blender, and significantly improved Windows usability for running CUDA applications on non-NVIDIA hardware.

2118

Mullvad VPN Controversy: Co-founder Daniel Berntsson's Political Donations

Mullvad VPN co-founder Daniel Berntsson's funding of the Swedish Örebro Party has sparked a debate over corporate neutrality and the intersection of private political donations and company values.

2119

Working With AI: A Concrete Example of the Sorcerer's Apprentice Problem

Carson Gross explores the strengths and weaknesses of AI in software development through a real-world bug fix in the hyperscript parser, highlighting the danger of blindly accepting AI-generated solutions.

2120

BAIR 2026 Graduate Showcase

The Berkeley Artificial Intelligence Research (BAIR) Lab announced its class of 2026 Ph.D. graduates, highlighting research across robotics, large language models, AI safety, and healthcare.

2121

genai-processors: a modular framework for building asynchronous and composable multimodal AI pipelines

A lightweight Python library for building modular, asynchronous, and composable AI pipelines that unify multimodal content processing and streaming.

2122

instill-core: an end-to-end AI infrastructure platform for data processing, pipeline orchestration, and model hosting

An end-to-end AI platform that simplifies the orchestration of unstructured data, AI pipelines, and model deployment to build RAG and AI-first applications.

2123

YTPro: a modified YouTube client with AI-powered video summarization and advanced playback controls

A modified YouTube client that integrates Google Gemini for AI video summarization alongside ad-blocking and content downloading tools.

2124

alan-sdk-web: an intelligent app platform SDK that enables real-time generation of business logic and UI

An SDK for embedding an intelligent layer into web applications that enables the real-time generation of business logic and UI components.

2125

presentation-ai: an open-source AI presentation generator with support for local LLMs and custom themes

An open-source, AI-powered presentation generator that creates customizable slides from a topic via an outline-first workflow.

2126

TTS-WebUI: a unified web interface for running and managing dozens of open-source text-to-speech and audio generation models

A unified web interface for managing and running a wide range of open-source text-to-speech, audio generation, and audio conversion AI models.

2127

Gemini-API: a reverse-engineered asynchronous Python wrapper for the Google Gemini web app

A reverse-engineered asynchronous Python wrapper for the Google Gemini web app that enables programmatic access to features like image generation, deep research, and custom Gems.

2128

hallucination-leaderboard: a public leaderboard tracking LLM hallucination rates in summarization tasks

A public leaderboard that uses Vectara's Hallucination Evaluation Model (HHEM) to measure and compare how often different LLMs hallucinate when summarizing documents.

2129

semantic-router: a superfast decision-making layer for LLMs and agents using semantic vector space for routing

A high-speed decision-making layer for LLMs and agents that uses semantic vector space to route requests based on meaning rather than slow LLM generation.

2130

transformer-explainer: an interactive browser-based visualization for learning the internal operations of GPT-2

An interactive visualization tool that runs a live GPT-2 model in the browser to help users learn how Transformer-based models predict text.

2131

ChatGPT-Shortcut: a curated prompt library and management tool for improving AI outputs across multiple platforms

An AI prompt management tool providing a curated library of 5,000+ prompts and tools to organize, create, and share custom prompts across various AI platforms.

2132

morphic: an AI search engine with a generative UI that renders rich inline components from streamed JSON

An AI-powered search engine that uses a generative UI to render rich, cited answers with interactive components instead of plain text.

2133

krita-ai-diffusion: a generative AI plugin for Krita that integrates diffusion models for precise image editing and painting

A Krita plugin that integrates generative AI diffusion models into the painting workflow, offering tools for inpainting, live painting, and precise structural control.

2134

outlines: a library for guaranteeing structured LLM outputs via type-constrained generation

A library for guaranteeing structured outputs from LLMs by constraining generation to match specific Python types or Pydantic models.

2135

Open-Generative-AI: an unrestricted open-source alternative to AI video platforms with local inference and multi-model support

An open-source, unrestricted AI studio for generating images and videos using 200+ models, featuring local inference options and a visual workflow builder.

2136

openui: an open-source AI-powered UI generator that renders live descriptions into framework-ready code

An open-source tool that lets you describe UI components in plain language and see them rendered live, with the ability to convert them to React, Svelte, or Web Components.

2137

TurboDiffusion: a video generation acceleration framework that reduces diffusion latency by 100-200x

TurboDiffusion is a video generation acceleration framework that speeds up diffusion generation by 100-200x on a single GPU using attention optimization and timestep distillation.

2138

MAGI-1: an autoregressive world model for scalable high-fidelity video generation with strong physical accuracy

MAGI-1 is an autoregressive video generation model that produces high-fidelity videos chunk-by-chunk to ensure temporal consistency and physical accuracy.

2139

ComfyUI-LTXVideo: custom ComfyUI nodes for advanced LTX-2 video generation and audio synthesis

A collection of custom ComfyUI nodes and workflows that extend the LTX-2 video generation model with features like HDR output, lip-syncing, and generative upscaling.

2140

transformerlab-app: an open-source machine learning platform that unifies AI research tooling and cluster orchestration

An open-source machine learning platform that unifies training, fine-tuning, inference, and evaluation into a single interface for individuals and research labs.

2141

What Happens When You Run a CUDA Kernel: From Source to SASS

An in-depth exploration of the CUDA execution pipeline, detailing how a simple vector addition kernel is compiled into SASS, launched via the NVIDIA driver, and executed across Streaming Multiprocessors.

2142

HackerRank Hiring Agent: The Risks of Non-Deterministic AI Resume Screening

An analysis of HackerRank's open-source hiring agent reveals significant scoring variance and design flaws, highlighting the dangers of using LLMs for high-stakes candidate evaluation.

2143

maestro: a streamlined tool to accelerate the fine-tuning of multimodal vision-language models

A streamlined tool for accelerating the fine-tuning of multimodal vision-language models like Florence-2, PaliGemma 2, and Qwen2.5-VL.

2144

SimpleTuner: a unified training framework for fine-tuning multi-modal generative models with enterprise-grade orchestration

A comprehensive training framework for fine-tuning image, video, and audio generative models, supporting a wide range of architectures with a focus on simplicity and memory efficiency.

2145

OneTrainer: a one-stop solution for training and fine-tuning a wide variety of diffusion models

A comprehensive training suite for diffusion models that provides tools for dataset preparation, fine-tuning, and model conversion through a GUI or CLI.

2146

OpenDeepWiki: an AI-driven repository knowledge base that generates structured docs, chat interfaces, and MCP endpoints from codebases

An AI-driven knowledge base generator that turns Git repositories and local directories into structured documentation, searchable chat interfaces, and MCP endpoints.

2147

MetaClaw: an agent proxy that enables AI assistants to meta-learn and evolve through real-world conversations

An agent proxy that enables AI assistants to meta-learn and evolve through real-world conversations using skill injection and asynchronous RL fine-tuning.

2148

Kiln: a local-first AI development workbench for prompt optimization, evaluations, and agent orchestration

A local-first AI development workbench that integrates prompt optimization, evaluations, RAG, and fine-tuning into a single workflow for teams.

2149

h2o-llmstudio: a no-code GUI and framework for fine-tuning large language models with support for memory-efficient training

A no-code GUI and framework for fine-tuning large language models, supporting memory-efficient techniques like LoRA and advanced optimization methods like DPO.

2150

CosyVoice: a scalable multilingual zero-shot text-to-speech synthesizer based on large language models

An LLM-based text-to-speech system for zero-shot multilingual speech synthesis with high speaker similarity and low-latency streaming.