AirSim: a high-fidelity visual and physical simulator for autonomous vehicles and AI research
An open-source simulator for drones and cars built on Unreal Engine, designed as a platform for AI research in deep learning, computer vision, and reinforcement learning.
MaaAssistantArknights: an image-recognition based automation assistant for Arknights that automates daily tasks and base management
An automation assistant for Arknights that uses image recognition and deep learning to automate daily tasks, base management, and resource farming.
vit-pytorch: a comprehensive collection of Vision Transformer (ViT) and its variants implemented in PyTorch
A comprehensive PyTorch library implementing the original Vision Transformer (ViT) and dozens of its advanced variants for image classification.
label-studio: a multi-modal open-source data labeling tool with ML-assisted pre-labeling and active learning
An open-source data labeling tool that allows users to annotate audio, text, images, and video to prepare or improve training data for machine learning models.
espnet: a comprehensive end-to-end speech processing toolkit for ASR, TTS, and spoken language understanding
An end-to-end speech processing toolkit that provides a unified framework for ASR, TTS, speech translation, and enhancement using PyTorch.
leon: an open-source personal AI assistant with agentic execution and local-first privacy
An open-source personal AI assistant that uses tools, memory, and agentic execution to perform tasks while supporting local AI models for privacy.
Cursor iOS App Installation Changes Account Privacy Settings
Installing the Cursor iOS app can irreversibly switch a user's account from Privacy Mode (Legacy) to a less restrictive Privacy Mode, removing the access to the previous setting.
PlayStation Store Removes Studio Canal Movies from User Libraries Without Refunds
Sony has removed hundreds of Studio Canal movies from PlayStation Store user accounts due to licensing agreements, sparking a debate over the legal nature of digital ownership.
US Supreme Court Rules Geofence Warrants Require Constitutional Protections
The US Supreme Court has ruled that geofence warrants must adhere to Fourth Amendment protections, rejecting the argument that opting into location services waives a user's privacy rights.
Tidal AI Policy: Demonetization and Labeling of AI-Generated Music
Tidal has introduced a new AI policy that demonetizes 100% AI-generated music and implements mandatory labeling to protect human artists and prevent fraudulent impersonations.
HunyuanVideo-1.5: a lightweight 8.3B parameter video generation model for high-quality synthesis on consumer GPUs
A lightweight 8.3B parameter video generation model that enables high-quality text-to-video and image-to-video synthesis on consumer-grade GPUs.
OpenMontage: an agentic video production system that orchestrates research, scripting, and editing into a full production pipeline
An open-source, agentic video production system that orchestrates research, scripting, asset generation, and editing to create professional videos from plain-language prompts.
Ornith-1.0 Release: Self-Improving Open-Source Models for Agentic Coding
Ornith-1.0 is a family of MIT-licensed, self-improving open-source models (9B to 397B) optimized for agentic coding and tool-calling, post-trained on Gemma 4 and Qwen 3.5.
Outer Shell: A Native Graphical Shell for SSH
Outer Shell introduces a browser-based graphical shell for remote servers that leverages SSH for security and Unix domain sockets for private, native app communication.
Rocket Lab Acquires Iridium
Rocket Lab has acquired Iridium, integrating a profitable satellite communications network to secure a baseline of regular launches and gain critical spectrum assets.
Samsung, SK Hynix, and Micron Sued in US for Memory Price Fixing
Samsung, SK Hynix, and Micron face a US lawsuit alleging they coordinated the reduction of DRAM supply and the discontinuation of older memory standards to inflate prices.
ZLUDA 6 Release: Running Unmodified CUDA Applications on Non-NVIDIA GPUs
ZLUDA 6 introduces support for 32-bit PhysX, basic texture support for Blender, and significantly improved Windows usability for running CUDA applications on non-NVIDIA hardware.
Mullvad VPN Controversy: Co-founder Daniel Berntsson's Political Donations
Mullvad VPN co-founder Daniel Berntsson's funding of the Swedish Örebro Party has sparked a debate over corporate neutrality and the intersection of private political donations and company values.
Working With AI: A Concrete Example of the Sorcerer's Apprentice Problem
Carson Gross explores the strengths and weaknesses of AI in software development through a real-world bug fix in the hyperscript parser, highlighting the danger of blindly accepting AI-generated solutions.
BAIR 2026 Graduate Showcase
The Berkeley Artificial Intelligence Research (BAIR) Lab announced its class of 2026 Ph.D. graduates, highlighting research across robotics, large language models, AI safety, and healthcare.
genai-processors: a modular framework for building asynchronous and composable multimodal AI pipelines
A lightweight Python library for building modular, asynchronous, and composable AI pipelines that unify multimodal content processing and streaming.
instill-core: an end-to-end AI infrastructure platform for data processing, pipeline orchestration, and model hosting
An end-to-end AI platform that simplifies the orchestration of unstructured data, AI pipelines, and model deployment to build RAG and AI-first applications.
YTPro: a modified YouTube client with AI-powered video summarization and advanced playback controls
A modified YouTube client that integrates Google Gemini for AI video summarization alongside ad-blocking and content downloading tools.
alan-sdk-web: an intelligent app platform SDK that enables real-time generation of business logic and UI
An SDK for embedding an intelligent layer into web applications that enables the real-time generation of business logic and UI components.
presentation-ai: an open-source AI presentation generator with support for local LLMs and custom themes
An open-source, AI-powered presentation generator that creates customizable slides from a topic via an outline-first workflow.
TTS-WebUI: a unified web interface for running and managing dozens of open-source text-to-speech and audio generation models
A unified web interface for managing and running a wide range of open-source text-to-speech, audio generation, and audio conversion AI models.
Gemini-API: a reverse-engineered asynchronous Python wrapper for the Google Gemini web app
A reverse-engineered asynchronous Python wrapper for the Google Gemini web app that enables programmatic access to features like image generation, deep research, and custom Gems.
hallucination-leaderboard: a public leaderboard tracking LLM hallucination rates in summarization tasks
A public leaderboard that uses Vectara's Hallucination Evaluation Model (HHEM) to measure and compare how often different LLMs hallucinate when summarizing documents.
semantic-router: a superfast decision-making layer for LLMs and agents using semantic vector space for routing
A high-speed decision-making layer for LLMs and agents that uses semantic vector space to route requests based on meaning rather than slow LLM generation.
transformer-explainer: an interactive browser-based visualization for learning the internal operations of GPT-2
An interactive visualization tool that runs a live GPT-2 model in the browser to help users learn how Transformer-based models predict text.
ChatGPT-Shortcut: a curated prompt library and management tool for improving AI outputs across multiple platforms
An AI prompt management tool providing a curated library of 5,000+ prompts and tools to organize, create, and share custom prompts across various AI platforms.
morphic: an AI search engine with a generative UI that renders rich inline components from streamed JSON
An AI-powered search engine that uses a generative UI to render rich, cited answers with interactive components instead of plain text.
krita-ai-diffusion: a generative AI plugin for Krita that integrates diffusion models for precise image editing and painting
A Krita plugin that integrates generative AI diffusion models into the painting workflow, offering tools for inpainting, live painting, and precise structural control.
outlines: a library for guaranteeing structured LLM outputs via type-constrained generation
A library for guaranteeing structured outputs from LLMs by constraining generation to match specific Python types or Pydantic models.
Open-Generative-AI: an unrestricted open-source alternative to AI video platforms with local inference and multi-model support
An open-source, unrestricted AI studio for generating images and videos using 200+ models, featuring local inference options and a visual workflow builder.
openui: an open-source AI-powered UI generator that renders live descriptions into framework-ready code
An open-source tool that lets you describe UI components in plain language and see them rendered live, with the ability to convert them to React, Svelte, or Web Components.
TurboDiffusion: a video generation acceleration framework that reduces diffusion latency by 100-200x
TurboDiffusion is a video generation acceleration framework that speeds up diffusion generation by 100-200x on a single GPU using attention optimization and timestep distillation.
MAGI-1: an autoregressive world model for scalable high-fidelity video generation with strong physical accuracy
MAGI-1 is an autoregressive video generation model that produces high-fidelity videos chunk-by-chunk to ensure temporal consistency and physical accuracy.
ComfyUI-LTXVideo: custom ComfyUI nodes for advanced LTX-2 video generation and audio synthesis
A collection of custom ComfyUI nodes and workflows that extend the LTX-2 video generation model with features like HDR output, lip-syncing, and generative upscaling.
transformerlab-app: an open-source machine learning platform that unifies AI research tooling and cluster orchestration
An open-source machine learning platform that unifies training, fine-tuning, inference, and evaluation into a single interface for individuals and research labs.
What Happens When You Run a CUDA Kernel: From Source to SASS
An in-depth exploration of the CUDA execution pipeline, detailing how a simple vector addition kernel is compiled into SASS, launched via the NVIDIA driver, and executed across Streaming Multiprocessors.
HackerRank Hiring Agent: The Risks of Non-Deterministic AI Resume Screening
An analysis of HackerRank's open-source hiring agent reveals significant scoring variance and design flaws, highlighting the dangers of using LLMs for high-stakes candidate evaluation.
maestro: a streamlined tool to accelerate the fine-tuning of multimodal vision-language models
A streamlined tool for accelerating the fine-tuning of multimodal vision-language models like Florence-2, PaliGemma 2, and Qwen2.5-VL.
SimpleTuner: a unified training framework for fine-tuning multi-modal generative models with enterprise-grade orchestration
A comprehensive training framework for fine-tuning image, video, and audio generative models, supporting a wide range of architectures with a focus on simplicity and memory efficiency.
OneTrainer: a one-stop solution for training and fine-tuning a wide variety of diffusion models
A comprehensive training suite for diffusion models that provides tools for dataset preparation, fine-tuning, and model conversion through a GUI or CLI.
OpenDeepWiki: an AI-driven repository knowledge base that generates structured docs, chat interfaces, and MCP endpoints from codebases
An AI-driven knowledge base generator that turns Git repositories and local directories into structured documentation, searchable chat interfaces, and MCP endpoints.
MetaClaw: an agent proxy that enables AI assistants to meta-learn and evolve through real-world conversations
An agent proxy that enables AI assistants to meta-learn and evolve through real-world conversations using skill injection and asynchronous RL fine-tuning.
Kiln: a local-first AI development workbench for prompt optimization, evaluations, and agent orchestration
A local-first AI development workbench that integrates prompt optimization, evaluations, RAG, and fine-tuning into a single workflow for teams.
h2o-llmstudio: a no-code GUI and framework for fine-tuning large language models with support for memory-efficient training
A no-code GUI and framework for fine-tuning large language models, supporting memory-efficient techniques like LoRA and advanced optimization methods like DPO.
CosyVoice: a scalable multilingual zero-shot text-to-speech synthesizer based on large language models
An LLM-based text-to-speech system for zero-shot multilingual speech synthesis with high speaker similarity and low-latency streaming.