latent-spaces/brag
A Claude Code skill that automatically transforms a software project into a short, shareable launch video with music and motion using Hyperframes.
TauricResearch/TradingAgents
TradingAgents is an open‑source, LangGraph‑based framework that orchestrates multiple LLM‑driven agents (fundamentals, sentiment, news, technical, researcher, trader, risk, portfolio) to simulate a full‑stack trading firm. It supports dozens of LLM providers, pulls point‑in‑time market data (SEC EDGAR, FRED, Yahoo Finance, social sentiment), offers checkpointed graph execution, logs decisions for continual learning, and provides a back‑testing CLI to evaluate performance across tickers and dates. Designed for research, not investment advice.
wide-trace/open-higgsfield
An open-source, self-hosted studio for generating images and videos using 32 different AI models via a unified prompt interface and user-provided API keys.
ApodexAI/FrontierAgent
FrontierAgent is an open‑source, terminal‑based runtime that lets you run autonomous ReAct agents or parallel Agent‑Team workflows. It provides a sandboxed filesystem, live TUI task board, asynchronous user interventions, and a built‑in benchmark harness for research‑grade long‑horizon tasks. Install with `uv`, point it at any OpenAI‑compatible model endpoint (including the free Apodex‑1.1 service), and run either `--mode react` for a single agent or `--mode agent_team` for coordinated parallel agents. The project is Apache‑2.0 licensed.
spinabot/brigade
Brigade is a self‑hosted, multi‑agent AI assistant framework. It lets you run a crew of agents that share persistent, provenance‑aware memory (Tideline), are organized in an org‑chart, and can be reached via terminal UI, WhatsApp, Slack, Discord, iMessage, etc. All API keys and data stay on your machine; you can bring any LLM (Claude, OpenAI, Gemini, local Ollama, etc.) and connect to 1 000+ third‑party apps through Composio. Features include skills, sub‑agent fan‑out, cron scheduling, document/media processing, optional Convex storage, and an MCP memory server. Install with a one‑line script or via npm, then onboard, start the always‑on gateway, and chat.
nexu-io/open-design
An open-source, local-first design platform that enables coding agents to generate and render web prototypes, decks, and motion graphics based on a brand's design system.
yi1108/printfilm
An open-source AI studio for creating short videos and episodic comics via a template-driven pipeline that automates storyboarding, image/video generation, and voiceovers.
Comfy-Org/ComfyUI
A modular, node-based AI engine for visual professionals to create complex workflows for generating images, video, audio, and 3D models with precise parameter control.
docling-project/docling
Docling is an open‑source Python library that parses a huge variety of document types (PDF, Word, slides, spreadsheets, e‑books, audio, video, emails, images, and many XML schemas) and enriches them with AI‑driven layout analysis, OCR, visual‑language‑model understanding, and speech‑to‑text. It outputs a unified `DoclingDocument` that can be exported to Markdown, HTML, JSON, DocLang, DocTags, etc., and integrates directly with LangChain, LlamaIndex, Crew AI, Haystack, or custom agents via an MCP or REST API. The tool runs locally for privacy‑sensitive data and also offers a CLI and service mode, making it ideal for building RAG pipelines, enterprise knowledge bases, financial or legal document processing, and multimodal AI agents.
higgsfield-ai/higgsfield
higgsfield is an open‑source GPU‑node manager and training framework for huge LLMs. It handles node allocation, ZeRO‑3/FSDP sharding, job queuing, and GitHub‑Actions‑based deployment, letting you train models like LLaMA‑70B across a cluster with a simple Python decorator.
milind-soni/OpenMausBot
OpenMausBot is a desktop chat application that lets you manage multiple AI agents (Claude, Codex, Grok, or any compatible CLI) as contacts. Each bot runs locally via a small harness server, can control a cloud or local computer, and can be connected to third‑party apps through Composio. The UI provides model switching, approval dialogs for risky actions, voice output via ElevenLabs, and a markdown‑based team import system. It runs on macOS, Windows, and Ubuntu, with all data stored locally unless you enable external services.
ruvnet/ruflo
Ruflo is an open‑source framework that adds a full agent execution layer around Claude Code/Codex. It provides 100+ specialized agents, swarm coordination, persistent vector memory, self‑learning loops, zero‑trust federation, a plugin marketplace, and a web UI for multi‑model chat. Install via `npx ruflo init` (full stack) or as Claude Code plugins (lite).
krillinai/OpenCreator
OpenCreator is an open‑source, locally‑run AI workspace that combines a Codex‑based conversational agent with a visual dashboard for multimodal content creation (video translation, downloading, thumbnail and image generation). It offers versioned workflows, local data storage, and a desktop client built on the same web UI. Supports many LLMs, Whisper, GPT‑Image, and other voice/translation services. Designed for creators and small teams who want AI‑assisted media production without relying on cloud‑only platforms.
ahujasid/mcp-for-blender
An MCP-based integration that connects LLMs to Blender, enabling prompt-driven 3D modeling, scene manipulation, and asset integration.
Anil-matcha/Open-Generative-AI
An unrestricted, open-source AI studio for generating images, videos, and audio using 400+ models without content filters or subscription fees.
gzxx-2025/aid-studio
A self-hosted AI production platform that integrates scriptwriting, storyboarding, image/video generation, and dubbing into a single workflow for creating AI movies, dramas, and comics.
microsoft/playwright-mcp
Playwright MCP is a Node.js server that lets LLM‑based agents control a Playwright browser via the Model Context Protocol. It returns structured accessibility snapshots (JSON) instead of screenshots, enabling token‑efficient, deterministic web automation with persistent browser state.
ZJU-REAL/Easel
Easel is an open‑source AI‑driven content workbench for social‑media creators. It combines an OpenClaw LLM agent, per‑account profiles, and 113+ executable “skills” to discover trends, plan topics, generate text, images, audio and video, publish directly to seven Chinese platforms, and feed performance data back into the profile for continual improvement.
lixiaoxiao9888-create/manju-laoli-skill
An industrial AIGC creation rulebase for AI agents that transforms scripts into consistent, model-ready video prompts using an Asset-First pipeline.
basketikun/infinite-canvas
An open-source AI image creation workbench that integrates an infinite canvas, multi-modal AI generation, and agentic canvas manipulation into a single workspace.
nilbuild/page-mascot
A React component and AI-powered toolset for adding interactive, cursor-tracking mascots to websites using sprite sheets.
OpenSenseNova/SenseNova-U1
OpenSenseNova’s SenseNova‑U1 series are open‑source, native‑unified multimodal models (text + image) built on the NEO‑unify architecture. They support high‑quality 4K image generation, dense‑text infographic creation, image editing, VQA, and interleaved text‑image tutorials. The repo provides training code, pre‑trained 8 B‑parameter checkpoints (including community‑quantized GGUF files), example inference scripts, and deployment guides, all under an Apache 2.0 license.
Devin-AXIS/deepseek-design
A native visual design system for DeepSeek Harness that allows AI to generate and manually refine editable websites, presentations, and videos.
chatfire-AI/huobao-drama
An AI-powered automated production platform that streamlines the creation of short dramas, from script generation and character design to final video synthesis.
HBAI-Ltd/Toonflow-app
An AI-powered workbench for short drama production that automates the workflow from scriptwriting and storyboarding to final video generation.
zenstory-ai/drama-skills
An AI-driven short-drama production workflow consisting of eleven agent skills that transform ideas or novels into scripts, storyboards, and video prompts with a focus on visual consistency.
alesha-pro/tools
A collection of local AI tools featuring optimized multi-GPU workflows for MiniMax H3 video-audio generation and specialized animation skills for AI agents.
v-modal/vmodal_sdk_flutter
A Flutter SDK for integrating multimodal semantic search across video and image libraries, enabling users to find moments by meaning, speech, or imagery.
omnigent-ai/omnigent
Omnigent is an open‑source meta‑harness that unifies many LLM coding agents (Claude Code, Codex, Cursor, etc.) under one CLI and web UI, letting you run, combine, and govern them across devices and cloud sandboxes, with built‑in collaboration and policy features.
darkzOGx/youtube-automation-agent
An open-source AI agent system that automates the entire YouTube channel lifecycle, from trend research and scriptwriting to video production, publishing, and analytics-driven optimization.
ningzimu/codex-ppt-skill
An AI agent skill that transforms articles, reports, and papers into image-based PowerPoint presentations by planning outlines, generating styled slides as images, and assembling them into a .pptx file.
EverettFish/holo-card-studio
A Codex Skill that transforms text or images into interactive 3D holographic cards with parallax effects, providing both a Three.js web viewer and an editable Blender project.
xuanyustudio/LocalMiniDrama
An open-source, local-first AI short drama creator that automates the entire production pipeline from script generation to final video synthesis using various AI APIs.
alchaincyf/huashu-design
A design-automation skill for AI agents that transforms text prompts into professional interactive prototypes, motion graphics, and editable presentation decks.
waooAI/waoowaoo
An AI creative workspace for images and videos that combines a brainstorming assistant and a visual canvas for organizing and refining AI-generated assets.
Merserk/dlss5-visual-enhancer
A Windows application for NVIDIA RTX GPUs that uses DLSS 5 and RTX Video technologies to upscale, enhance, and interpolate frames for images, videos, and live streams.
Open-LLM-VTuber/Open-LLM-VTuber
An open-source, voice-interactive AI companion featuring a Live2D avatar and visual perception, capable of running entirely offline for private, real-time conversations.
Swayingleaves/novanova-studio
An AI Agent-driven visual creation workbench that integrates image and video generation into an infinite canvas to maintain creative context.
yejy53/Editable-Design
Editable Visual Design is an open‑source toolkit that lets an LLM (via Codex “Skills”) turn natural‑language design briefs into fully editable graphics—HTML pages or PowerPoint files with independent text, image, and layout layers. It ships three skills: `paper-fig` for research diagrams, `editable-design` for posters/infographics, and `html-to-pptx` for converting HTML to PPTX. The system records the LLM’s design steps (Agent Design Replay) and provides a visual editor, enabling rapid AI‑generated drafts that remain editable locally.
ningzimu/image-to-editable-ppt-skill
A multi-agent AI skill that converts images, PDFs, and image-based slides into editable PowerPoint presentations by reconstructing text, shapes, and assets.
QwenLM/Qwen-MM-Plugins
A collection of multimodal plugins for Qwen models that enable AI agents to natively process images, video, audio, and 3D files, and control software like Blender and FreeCAD.
techjarves/Portable-Local-Studio
A zero-configuration, offline AI studio that integrates Stable Diffusion, LLMs, Whisper, and Kokoro TTS into a single hardware-accelerated desktop interface.
bytedance/UI-TARS-desktop
A multimodal AI agent stack that enables natural language control of desktop applications and web browsers through visual recognition and precise GUI interaction.
buxuku/SmartSub
SmartSub (妙幕) is an open‑source, cross‑platform desktop app that lets you download a video, run local or cloud speech‑to‑text, translate subtitles with many services, edit/proofread them, synthesize voice (including zero‑shot voice cloning), and finally hard‑burn or mux the subtitles back into the video. It runs offline with optional GPU acceleration and can be used completely for free.
xpzouying/xiaohongshu-mcp
A self‑hosted MCP server that lets AI assistants log in to Xiaohongshu, publish posts (images or video), search feeds, like/favorite, comment, and fetch user or post details. Available as binaries, Docker image, or source build, and integrates with Claude Code, Cursor, VS Code, Open Code, Cline, etc.
PurpleDoubleD/locally-uncensored
A comprehensive local AI studio for desktop that integrates chat, image and video generation, and a coding agent into a single, easy-to-install application.
lightningpixel/modly
A local, open-source desktop application that turns photos into 3D meshes using AI models running on the user's GPU.
Tencent/WeMM-Embedding
WeMM‑Embedding is Tencent’s open‑source multimodal embedding suite (2B/4B/9B parameters) that converts text, images, video, and visual documents into a single L2‑normalised vector. It supports a range of output dimensions via “Matryoshka” truncation, offers ready‑made inference scripts for 🤗 Transformers and Sentence‑Transformers, and includes serving wrappers for vLLM and SGLang. Benchmarks (MMEB‑v2/v3) show state‑of‑the‑art scores across image, video, document, text, and agent tasks. Models are hosted on Hugging Face and released under Apache 2.0.