microsoft/markitdown
Python tool for converting files and office documents to Markdown.
66 starred repositories in this area.
66 result(s)
Python tool for converting files and office documents to Markdown.
The context API to search, scrape, and interact with the web at scale. 🔥
The Memory Layer for AI Agents - Drop-in memory infrastructure for AI agents and apps. Context that persists. Built for production.
LlamaIndex is the leading document agent and OCR platform
Build resilient agents.
[EMNLP2025] LightRAG: Simple and Fast Retrieval-Augmented Generation
A modular graph-based Retrieval-Augmented Generation (RAG) system
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
Build Real-Time Knowledge Graphs for AI Agents
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
Memory and context engine + app that is extremely fast, scalable, and can be run fully locally. The Memory API for the AI era.
"RAG-Anything: All-in-One RAG Framework"
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
How Python does AI. Agents, realtime voice, image generation, embeddings. Every model, every interface, typed end to end.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
Pocket Flow: 100-line LLM framework. Let Agents build Agents!
X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.
"Context engineering is the delicate art and science of filling the context window with just the right information for the next step." — Andrej Karpathy. A frontier, first-principles handbook inspired by Karpathy and 3Blue1Brown for moving beyond prompt engineering to the wider discipline of context design, orchestration, and optimization.
PraisonAI 🦞 — Hire a 24/7 AI Workforce. Stop writing boilerplate and start shipping autonomous self-improving agents that research, plan, code, and execute tasks. Deployed in 5 lines of code with built-in memory, RAG, and support for 100+ LLMs.
🧑🚀 全世界最好的LLM资料总结(多模态生成、Agent、辅助编程、AI审稿、数据处理、模型训练、模型推理、o1 模型、MCP、小语言模型、视觉语言模型) | Summary of the world's best LLM resources.
Large Action Model framework to develop AI Web Agents
Main reference implementation for NLWeb, implemented in Python.
250+ Fine-tuning & RL Notebooks for text, vision, audio, embedding, TTS models.
Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL
Knowledge Agents and Management in the Cloud
MongoDB's Generative AI Showcase: an exhaustive collection of examples and sample applications covering Retrieval-Augmented Generation (RAG), AI agents, and industry-specific use cases.
Everything you need to know to build your own RAG application
A simple, easy-to-hack GraphRAG implementation
⚡FlashRAG: A Python Toolkit for Efficient RAG Research (WWW2025 Resource)
ReMe: Memory Management Kit for Agents - Remember Me, Refine Me.
Learn to build your Second Brain AI assistant with LLMs, agents, RAG, fine-tuning, LLMOps and AI systems techniques.
AI reads books: Page-by-Page PDF Knowledge Extractor & Summarizer. script performs an intelligent page-by-page analysis of PDF books, methodically extracting knowledge points and generating progressive summaries at specified intervals
NeMo Retriever Library is a scalable, performance-oriented document content and metadata extraction microservice. NeMo Retriever Library uses specialized NVIDIA NIM microservices to find, contextualize, and extract text, tables, charts and images that you can use in downstream generative applications.
User Profile-Based Long-Term Memory for AI Chatbot Applications.
Awesome papers about unifying LLMs and KGs
PIKE-RAG: sPecIalized KnowledgE and Rationale Augmented Generation
[Paper List] Papers integrating knowledge graphs (KGs) and large language models (LLMs)
Perplexity style AI Search engine clone built with Gemini 2.0 Flash and Grounding
[ACL2026] "MiniRAG: Making RAG Simpler with Small and Open-Sourced Language Models"
ContextGem: Effortless LLM extraction from documents
WFGY is heading toward WFGY 5.0 Polaris Protocol, a major open-source release for AI reasoning, RAG, agents, and real-world workflows. Includes Problem Map, Global Debug Card, WFGY 4.0, and the CFV Easter Egg.
Optimized Agentic and LLM Bulk Processing Over Your Data
the resources about the application based on LLM with RAG pattern
This repository includes the official implementation of OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs.
Comprehensive guide to learn RAG from basics to advanced.
Awesome-LLM-RAG: a curated list of advanced retrieval augmented generation (RAG) in Large Language Models
Daily updated LLM papers. 每日更新 LLM 相关的论文,欢迎订阅 👏 喜欢的话动动你的小手 🌟 一个
[NeurIPS '25] Knowledge Graph Generation from Any Text
utilities for decoding deep representations (like sentence embeddings) back to text
Local models support for Microsoft's graphrag using ollama (llama3, mistral, gemma2 phi3)- LLM & Embedding extraction
Empower Large Language Models (LLM) using Knowledge Graph based Retrieval-Augmented Generation (KG-RAG) for knowledge intensive tasks
Microsoft's GraphRAG + AutoGen + Ollama + Chainlit = Fully Local & Free Multi-Agent RAG Superbot
A Graph RAG System for Evidenced-based Medical Information Retrieval [ACL 2025]
Evidence-first local memory for AI agents with temporal versions, admission policies, citations, explainable recall, MCP, and audit tooling.
👩🏻🍳 A collection of example notebooks using Haystack
[EMNLP 2025] OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking
Parse PDFs into markdown using Vision LLMs
RAG 相關資源、實作與評估方法的精選清單。
日常文件轉 Markdown —— 涵蓋 PDF、Office、Apple Keynote/Numbers、EPUB 電子書等 16 種格式。中文友好、表格保留、隱私優先、全程本地。
醫療領域的檢索增強生成研究專案。
FlexRAG: A RAG Framework for Information Retrieval and Generation.
Build Agents That Recall What Matters. Systematically engineer relevant context from chat history & business data. (Python Client)
The official repository for the paper: Evaluation of Retrieval-Augmented Generation: A Survey.
用知識圖譜替 LLM 的回答找事實依據的實驗性專案。
RAG-powered documentation assistant that converts, processes, and provides semantic search capabilities for Odoo's technical documentation. Supports multiple Odoo versions with an interactive chat interface powered by LLM models.
用知識圖譜做 RAG 的教學 notebook。