Speech, Vision & Media

78 starred repositories in this area.

78 result(s)

Comfy-Org/ComfyUI

The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.

★ 132.1k · Python · updated 2026-09

Pythonaicomfycomfyui

harry0703/MoneyPrinterTurbo

利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.

★ 121.5k · Python · updated 2026-09

Pythonai-video-generatorcontent-creationffmpeg

RVC-Boss/GPT-SoVITS

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

★ 61.7k · Python · updated 2026-08

Pythontext-to-speechttsvits

remotion-dev/remotion

🎥 Make videos programmatically with React

★ 58.6k · TypeScript · updated 2026-09

TypeScriptjavascriptreactvideo

microsoft/VibeVoice

Open-Source Frontier Voice AI

★ 54k · Python · updated 2026-09

Python

suno-ai/bark

🔊 Text-Prompted Generative Audio Model

★ 39.3k · Jupyter Notebook · updated 2024-08

Jupyter Notebook

OpenBMB/VoxCPM

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

★ 36.9k · Python · updated 2026-09

Pythonaudiodeeplearningminicpm

Zackriya-Solutions/meetily

Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (Meetly Ai - https://meetily.ai) is the #1 Self-hosted, Open-source Ai meeting note taker for macOS & Windows. Understand How to write meeting minutes

★ 30.5k · Rust · updated 2026-09

Rustaiai-meeting-assistantllm

hpcaitech/Open-Sora

Open-Sora: Democratizing Efficient Video Production for All

★ 29.7k · Python · updated 2026-04

Python

ATH-MaaS/Pixelle-Video

🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine

★ 27.9k · Python · updated 2026-06

Pythonaigccomfyuiimage-generation

resemble-ai/chatterbox

SoTA open-source TTS

★ 26.3k · Python · updated 2026-07

Python

SYSTRAN/faster-whisper

Faster Whisper transcription with CTranslate2

★ 25.3k · Python · updated 2025-11

Pythondeep-learninginferenceopenai

QwenLM/Qwen3-VL

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

★ 19.9k · Jupyter Notebook · updated 2026-01

Jupyter Notebook

facebookresearch/sam2

The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.

★ 19.8k · Jupyter Notebook · updated 2026-05

Jupyter Notebook

nari-labs/dia

A TTS model capable of generating ultra-realistic dialogue in one pass.

★ 19.4k · Python · updated 2025-11

Pythonaiopen-weighttext-to-speech

KlingAIResearch/LivePortrait

Bring portraits to life!

★ 19k · Python · updated 2026-06

Pythonface-animationimage-animationvideo-editing

BradyFU/Awesome-Multimodal-Large-Language-Models

✨✨Latest Advances on Multimodal Large Language Models

★ 18k · — · updated 2026-09

chain-of-thoughtin-context-learninginstruction-following

deepseek-ai/Janus

Janus-Series: Unified Multimodal Understanding and Generation Models

★ 17.8k · Python · updated 2025-02

Pythonany-to-anyfoundation-modelsllm

Wan-Video/Wan2.1

Wan: Open and Advanced Large-Scale Video Generative Models

★ 16.9k · Python · updated 2026-03

Pythonaigcvideogeneration

datawhalechina/leedl-tutorial

《李宏毅深度学习教程》(李宏毅老师推荐👍,苹果书🍎),PDF下载地址:https://github.com/datawhalechina/leedl-tutorial/releases

★ 16.8k · Jupyter Notebook · updated 2025-11

Jupyter Notebookbertchatgptcnn

SWivid/F5-TTS

Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"

★ 15.2k · Python · updated 2026-07

Python

duixcom/Duix-Avatar

🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.

★ 15.1k · C · updated 2026-04

Cai-avatarai-avatarscloning

microsoft/TRELLIS

Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).

★ 13.6k · Python · updated 2026-06

Python3d3d-aigc3d-generation

Vaibhavs10/insanely-fast-whisper

把 Whisper 榨到極快的推論腳本,幾分鐘轉完數小時音檔。

★ 13.1k · Jupyter Notebook · updated 2025-10

Jupyter Notebook

ace-step/ACE-Step-1.5

The most powerful local music generation model that outperforms almost all commercial alternatives, supporting Mac, AMD, Intel, and CUDA devices.

★ 12.6k · Python · updated 2026-09

Pythontext2music

rany2/edge-tts

Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key

★ 11.9k · Python · updated 2026-03

Pythonspeech-synthesistext-to-speechtts

qubvel-org/segmentation_models.pytorch

Semantic segmentation models with 500+ pretrained convolutional and transformer-based backbones.

★ 11.7k · Python · updated 2026-09

Pythoncomputer-visiondeeplab-v3-plusdeeplabv3

QuentinFuxa/WhisperLiveKit

Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.

★ 11k · Python · updated 2026-09

Pythonautomatic-speech-recognitionfastapipython

SparkAudio/Spark-TTS

Spark-TTS Inference Code

★ 11k · Python · updated 2025-04

Python

KoljaB/RealtimeSTT

A robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.

★ 10.1k · Python · updated 2026-08

Pythonpythonrealtimespeech-to-text

PeterL1n/RobustVideoMatting

Robust Video Matting in PyTorch, TensorFlow, TensorFlow.js, ONNX, CoreML!

★ 9.5k · Python · updated 2024-04

Pythonaicomputer-visiondeep-learning

deepbeepmeep/Wan2GP

A fast AI Video Generator for the GPU Poor. Supports Wan 2.1/2.2, LTX-2, Qwen Image, Hunyuan Video, LTX Video and Flux.

★ 9.2k · Python · updated 2026-09

Pythonaifluxflux2

studio-dots-ai/dots.ocr

Multilingual Document Layout Parsing in a Single Vision-Language Model

★ 9.1k · Python · updated 2026-03

Python

jianchang512/clone-voice

A sound cloning tool with a web interface, using your voice or any sound to record audio / 一个带web界面的声音克隆工具,使用你的音色或任意声音来录制音频

★ 9k · Python · updated 2025-08

Pythonclonevoicespeech-analysisstsarchived

JimmyLv/awesome-nano-banana

Awesome curated collection of images and prompts generated by gemini-2.5-flash-image (aka Nano Banana) state-of-the-art image generation and editing model. Explore AI generated visuals created with Gemini, showcasing Google’s advanced image generation capabilities.

★ 8.8k · JavaScript · updated 2025-09

JavaScriptchatgptflux-kontextgemini-2-5-flash-image

boson-ai/higgs-audio

Text-audio foundation model from Boson AI

★ 8.3k · Python · updated 2026-06

Python

jamez-bondos/awesome-gpt4o-images

Awesome curated collection of images and prompts generated by GPT-4o and gpt-image-1. Explore AI generated visuals created with ChatGPT and Sora, showcasing OpenAI’s advanced image generation capabilities.

★ 8.1k · JavaScript · updated 2025-05

JavaScriptai-artai-image-examplesanime-ai-art

Zyphra/Zonos

Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers.

★ 7.2k · Python · updated 2025-03

Python

ngxson/smolvlm-realtime-webcam

Real-time webcam demo with SmolVLM and llama.cpp server

★ 5.6k · HTML · updated 2025-05

HTML

modstart-lib/aigcpanel

AIGCPanel 是一个简单易用的一站式AI数字人系统,支持视频合成、声音合成、声音克隆,简化本地模型管理、一键导入和使用AI模型。

★ 5.5k · TypeScript · updated 2026-09

TypeScriptaiaigccosyvoice

OHF-Voice/piper1-gpl

Fast and local neural text-to-speech engine

★ 5.5k · C++ · updated 2026-09

C++

nICEnnnnnnnLee/BilibiliDown

(GUI-多平台支持) B站 哔哩哔哩 视频下载器。支持稍后再看、收藏夹、UP主视频批量下载|Bilibili Video Downloader 😳

★ 5.3k · Java · updated 2026-07

Javabilibilicookiedownload-videos

Anil-matcha/AI-Youtube-Shorts-Generator

Open-source alternative to Opus Clip, Vidyo.ai, Klap & SubMagic. Turn long-form YouTube videos into viral 9:16 shorts using LLM highlight detection, Whisper transcription, and auto vertical cropping — free, no watermarks, no per-clip credits.

★ 4.9k · Python · updated 2026-09

Python2short-ai-alternativeai-clip-generatorai-clipping

SandAI-org/MAGI-1

MAGI-1: Autoregressive Video Generation at Scale

★ 3.8k · Python · updated 2026-06

Pythonautoregressivediffusion-modelsvideo-generation

NExT-GPT/NExT-GPT

Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model

★ 3.6k · Python · updated 2025-05

Pythonchatgptfoundation-modelsgpt-4

Soul-AILab/SoulX-Podcast

SoulX-Podcast is an inference codebase by the Soul AI team for generating high-fidelity podcasts from text.

★ 3.5k · Python · updated 2025-12

Python

QwenLM/Qwen3-ASR

Qwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilingual speech/music/song recognition, language detection and timestamp prediction.

★ 3.5k · Python · updated 2026-06

Python

MiniMax-AI/MiniMax-01

The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention

★ 3.5k · Python · updated 2025-07

Pythonlarge-language-modelsllmllms

chenyme/Chenyme-AAVT

这是一个全自动(音频)视频翻译项目。利用Whisper识别声音,AI大模型翻译字幕,最后合并字幕视频,生成翻译后的视频。

★ 3.1k · Python · updated 2025-04

Pythonfaster-whispergpt-4gpt-4oarchived

axcore/tartube

A GUI front-end for youtube-dl, partly based on youtube-dl-gui and written in Python 3 / Gtk 3

★ 3.1k · Python · updated 2026-07

Python

LeeJunHyun/Image_Segmentation

Pytorch implementation of U-Net, R2U-Net, Attention U-Net, and Attention R2U-Net.

★ 3.1k · Python · updated 2023-06

Python

Tramac/awesome-semantic-segmentation-pytorch

Semantic Segmentation on PyTorch (include FCN, PSPNet, Deeplabv3, Deeplabv3+, DANet, DenseASPP, BiSeNet, EncNet, DUNet, ICNet, ENet, OCNet, CCNet, PSANet, CGNet, ESPNet, LEDNet, DFANet)

★ 3.1k · Python · updated 2023-01

Pythonpytorchsemantic-segmentation

bytedance/InfiniteYou

🔥 [ICCV 2025 Highlight] InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity

★ 2.7k · Python · updated 2025-08

Pythondiffusersdiffusiondiffusion-transformer

miantiao-me/hacker-podcast

一个基于 AI 的 Hacker News 中文播客项目,每天自动抓取 Hacker News 热门文章,通过 AI 生成中文总结并转换为播客内容。

★ 2.6k · TypeScript · updated 2026-09

TypeScriptaiai-agentai-workflow

alecm20/story-flicks

使用AI大模型,一键生成高清故事短视频。Generate high-definition story short videos with one click using AI large models.

★ 2.5k · Python · updated 2025-03

Pythonai-videoai-video-generatorchatgpt

harry0703/AudioNotes

快速提取音视频内容,整理成一份结构化的markdown笔记

★ 2.5k · Python · updated 2026-08

Pythonaiasrfunasr

PKU-YuanGroup/LLaVA-CoT

[ICCV 2025] LLaVA-CoT, a visual language model capable of spontaneous, systematic reasoning

★ 2.1k · Python · updated 2025-12

Python

QwenLM/Qwen2-Audio

The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.

★ 2.1k · Python · updated 2025-04

Python

Hao0321/video-autopilot-kit

Fill-in-your-own-data framework for YouTube / short-form video automation: CapCut JSON + ffmpeg tooling + an onboarding questionnaire. Ships with zero private data.

★ 2.1k · Python · updated 2026-08

Pythoncapcutcontent-creationcreator-tools

wwbin2017/bailing

百聆 是一个类似GPT-4o的语音对话机器人,通过ASR+LLM+TTS实现,集成DeepSeek R1等优秀大模型,接入openClaw,真正的个人语音助手,时延低至800ms,Mac等低配置也可运行,支持打断

★ 1.8k · Python · updated 2026-04

Pythonaiasrchatgpt

FoloUp/FoloUp

💼 Your own AI-powered voice interviewer for hiring.

★ 1.3k · TypeScript · updated 2026-05

TypeScripthiringinterviewingrecruiting

aTrainTranscription/aTrain

A GUI tool for offline transcription of speech recordings, including speaker diarization, utilizing state-of-the-art machine learning models.

★ 1.2k · Python · updated 2026-09

Python

senstella/parakeet-mlx

An implementation of the Nvidia's Parakeet models for Apple Silicon using MLX.

★ 977 · Python · updated 2026-08

Python

ImprintLab/Medical-SAM2

Medical SAM 2: Segment 3D Medical Images Via Segment Anything Model 2

★ 930 · Python · updated 2025-01

Pythondeep-learningmedicalmedical-imaging

DataoceanAI/Dolphin

Dolphin is a multilingual, multitask ASR model jointly trained by DataoceanAI and Tsinghua University.

★ 787 · Python · updated 2026-06

Python

imelnyk/ArxivPapers

Code behind Arxiv Papers

★ 544 · Python · updated 2024-04

Python

chrischoy/WhisperChain

Speech to Text but with all the bells and whistles and most importantly AI! AI will clean up your filler words, edit and will refine what you said!

★ 332 · Python · updated 2025-02

Python

AS-AIGC/TranscriptHub

逐字稿處理平台,把語音內容集中管理與轉錄。

★ 245 · JavaScript · updated 2026-06

JavaScript

PCL-Voice/PengChengStarling

PengChengStarling is specifically designed for developing multilingual ASR models based on the icefall project, supporting a complete ASR pipeline that includes data processing, model training, inference, fine-tuning, and deployment.

★ 190 · Python · updated 2025-03

Python

thepersonalaicompany/amurex-backend

This is the backend for the entire Amurex project.

★ 147 · Python · updated 2025-04

Python

hanifabd/voice-activity-detection-vad-realtime

Real-time Voice Activity Detection (VAD) with some example use case like simple voice bot and live transcription (realtime transcription)

★ 114 · Python · updated 2025-08

Pythonlive-transcriptmachine-learningrealtime-transcribe

deepbeepmeep/Cosmos1GP

Cosmos1GP for the GPU Poor by DeepBeepMeep

★ 92 · Python · updated 2025-02

Python

mtkresearch/generative-fusion-decoding

Generative Fusion Decoding (GFD) is a novel framework for integrating Large Language Models (LLMs) into multi-modal text recognition systems like ASR and OCR, improving performance and efficiency by enabling seamless fusion without requiring re-training.

★ 87 · Python · updated 2025-07

Python

Winston774/ai-music-channel-starter

AI 音樂 YouTube 頻道自動化 starter repo:生成、審查、上傳與數據追蹤

★ 53 · TypeScript · updated 2026-05

TypeScript

Farzad-R/Finetune-LLAVA-NEXT

This repository contains codes for fine-tuning LLAVA-1.6-7b-mistral (Multimodal LLM) model.

★ 41 · Jupyter Notebook · updated 2024-11

Jupyter Notebook

HighCWu/flux-4bit

把 FLUX 生圖模型量化到 4bit,讓小顯存也能跑。

★ 23 · Jupyter Notebook · updated 2024-08

Jupyter Notebook

myyang19770915/MOSS-ASR

Stream ASR and adding LLM 智慧比對校稿

★ 20 · Python · updated 2026-09

Python

ChihchengHsieh/Multimodal-Medical-Diagnosis-System

多模態醫療診斷系統的研究實作。

★ 11 · Jupyter Notebook · updated 2022-03

Jupyter Notebook

中文版