Comfy-Org/ComfyUI
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
78 starred repositories in this area.
78 result(s)
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
🎥 Make videos programmatically with React
Open-Source Frontier Voice AI
🔊 Text-Prompted Generative Audio Model
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (Meetly Ai - https://meetily.ai) is the #1 Self-hosted, Open-source Ai meeting note taker for macOS & Windows. Understand How to write meeting minutes
Open-Sora: Democratizing Efficient Video Production for All
🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine
SoTA open-source TTS
Faster Whisper transcription with CTranslate2
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
A TTS model capable of generating ultra-realistic dialogue in one pass.
Bring portraits to life!
✨✨Latest Advances on Multimodal Large Language Models
Janus-Series: Unified Multimodal Understanding and Generation Models
Wan: Open and Advanced Large-Scale Video Generative Models
《李宏毅深度学习教程》(李宏毅老师推荐👍,苹果书🍎),PDF下载地址:https://github.com/datawhalechina/leedl-tutorial/releases
Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.
Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).
把 Whisper 榨到極快的推論腳本,幾分鐘轉完數小時音檔。
The most powerful local music generation model that outperforms almost all commercial alternatives, supporting Mac, AMD, Intel, and CUDA devices.
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
Semantic segmentation models with 500+ pretrained convolutional and transformer-based backbones.
Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.
Spark-TTS Inference Code
A robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.
Robust Video Matting in PyTorch, TensorFlow, TensorFlow.js, ONNX, CoreML!
A fast AI Video Generator for the GPU Poor. Supports Wan 2.1/2.2, LTX-2, Qwen Image, Hunyuan Video, LTX Video and Flux.
Multilingual Document Layout Parsing in a Single Vision-Language Model
A sound cloning tool with a web interface, using your voice or any sound to record audio / 一个带web界面的声音克隆工具,使用你的音色或任意声音来录制音频
Awesome curated collection of images and prompts generated by gemini-2.5-flash-image (aka Nano Banana) state-of-the-art image generation and editing model. Explore AI generated visuals created with Gemini, showcasing Google’s advanced image generation capabilities.
Text-audio foundation model from Boson AI
Awesome curated collection of images and prompts generated by GPT-4o and gpt-image-1. Explore AI generated visuals created with ChatGPT and Sora, showcasing OpenAI’s advanced image generation capabilities.
Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers.
Real-time webcam demo with SmolVLM and llama.cpp server
AIGCPanel 是一个简单易用的一站式AI数字人系统,支持视频合成、声音合成、声音克隆,简化本地模型管理、一键导入和使用AI模型。
Fast and local neural text-to-speech engine
(GUI-多平台支持) B站 哔哩哔哩 视频下载器。支持稍后再看、收藏夹、UP主视频批量下载|Bilibili Video Downloader 😳
Open-source alternative to Opus Clip, Vidyo.ai, Klap & SubMagic. Turn long-form YouTube videos into viral 9:16 shorts using LLM highlight detection, Whisper transcription, and auto vertical cropping — free, no watermarks, no per-clip credits.
MAGI-1: Autoregressive Video Generation at Scale
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
SoulX-Podcast is an inference codebase by the Soul AI team for generating high-fidelity podcasts from text.
Qwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilingual speech/music/song recognition, language detection and timestamp prediction.
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention
这是一个全自动(音频)视频翻译项目。利用Whisper识别声音,AI大模型翻译字幕,最后合并字幕视频,生成翻译后的视频。
A GUI front-end for youtube-dl, partly based on youtube-dl-gui and written in Python 3 / Gtk 3
Pytorch implementation of U-Net, R2U-Net, Attention U-Net, and Attention R2U-Net.
Semantic Segmentation on PyTorch (include FCN, PSPNet, Deeplabv3, Deeplabv3+, DANet, DenseASPP, BiSeNet, EncNet, DUNet, ICNet, ENet, OCNet, CCNet, PSANet, CGNet, ESPNet, LEDNet, DFANet)
🔥 [ICCV 2025 Highlight] InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
一个基于 AI 的 Hacker News 中文播客项目,每天自动抓取 Hacker News 热门文章,通过 AI 生成中文总结并转换为播客内容。
使用AI大模型,一键生成高清故事短视频。Generate high-definition story short videos with one click using AI large models.
快速提取音视频内容,整理成一份结构化的markdown笔记
[ICCV 2025] LLaVA-CoT, a visual language model capable of spontaneous, systematic reasoning
The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.
Fill-in-your-own-data framework for YouTube / short-form video automation: CapCut JSON + ffmpeg tooling + an onboarding questionnaire. Ships with zero private data.
百聆 是一个类似GPT-4o的语音对话机器人,通过ASR+LLM+TTS实现,集成DeepSeek R1等优秀大模型,接入openClaw,真正的个人语音助手,时延低至800ms,Mac等低配置也可运行,支持打断
💼 Your own AI-powered voice interviewer for hiring.
A GUI tool for offline transcription of speech recordings, including speaker diarization, utilizing state-of-the-art machine learning models.
An implementation of the Nvidia's Parakeet models for Apple Silicon using MLX.
Medical SAM 2: Segment 3D Medical Images Via Segment Anything Model 2
Dolphin is a multilingual, multitask ASR model jointly trained by DataoceanAI and Tsinghua University.
Code behind Arxiv Papers
Speech to Text but with all the bells and whistles and most importantly AI! AI will clean up your filler words, edit and will refine what you said!
逐字稿處理平台,把語音內容集中管理與轉錄。
PengChengStarling is specifically designed for developing multilingual ASR models based on the icefall project, supporting a complete ASR pipeline that includes data processing, model training, inference, fine-tuning, and deployment.
This is the backend for the entire Amurex project.
Real-time Voice Activity Detection (VAD) with some example use case like simple voice bot and live transcription (realtime transcription)
Cosmos1GP for the GPU Poor by DeepBeepMeep
Generative Fusion Decoding (GFD) is a novel framework for integrating Large Language Models (LLMs) into multi-modal text recognition systems like ASR and OCR, improving performance and efficiency by enabling seamless fusion without requiring re-training.
AI 音樂 YouTube 頻道自動化 starter repo:生成、審查、上傳與數據追蹤
This repository contains codes for fine-tuning LLAVA-1.6-7b-mistral (Multimodal LLM) model.
把 FLUX 生圖模型量化到 4bit,讓小顯存也能跑。
Stream ASR and adding LLM 智慧比對校稿
多模態醫療診斷系統的研究實作。