Back to Rankings返回排行榜
Top 100 · Multimodal AI前 100 · 多模态 AI
100 repositories sorted by multimodal ai 按 多模态 AI 排序,共 100 个仓库
| # | Repository仓库 | Stars | Forks | Language语言 | Issues | Description描述 | Last Commit最后提交 |
|---|---|---|---|---|---|---|---|
| 1 | transformers huggingface | 167.3k | 34.8k | Python | 712 | 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training. 🤗 Transformers:文本、视觉、音频和多模态模型中最先进的机器学习模型的模型定义框架,用于推理和训练。 | 2026-10-11 |
| 2 | anything-llm Mintplex-Labs | 66.9k | 7.5k | JavaScript | 309 | Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience 别再出租你的智力了。 AnythingLLM 拥有它。强大的本地优先代理体验所需的一切 | 2026-10-09 |
| 3 | ai-agent-book bojieli | 53.4k | 6.0k | Python | 12 | 《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码《深入理解 AI Agent:设计原理与工程实践》(李博杰 着)开源主仓库:全书正文、编译版 PDF 与按章配套代码 | 2026-10-08 |
| 4 | UI-TARS-desktop bytedance | 39.2k | 4.0k | TypeScript | 349 | The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra开源多模式 AI 代理堆栈:连接尖端 AI 模型和代理基础设施 | 2026-10-09 |
| 5 | sglang sgl-project | 37.0k | 9.4k | Python | 930 | SGLang is a high-performance serving framework for large language models and multimodal models.SGLang 是一个用于大型语言模型和多模态模型的高性能服务框架。 | 2026-10-11 |
| 6 | haystack deepset-ai | 26.7k | 3.3k | Python | 101 | Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search, and conversational systems.开源 AI 编排框架,用于构建上下文工程、生产就绪的 LLM 应用程序。通过对检索、路由、内存和生成的显式控制来设计模块化管道和代理工作流程。专为可扩展代理、RAG、多模式应用程序、语义搜索和对话系统而构建。 | 2026-10-10 |
| 7 | LLaVA haotian-liu | 25.1k | 2.8k | Python | 1098 | [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.[NeurIPS'23 Oral] 视觉指令调优 (LLaVA) 旨在实现 GPT-4V 级别及以上的功能。 | 2024-08-12 |
| 8 | unilm microsoft | 22.2k | 2.7k | Python | 648 | Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities跨任务、语言和模式的大规模自监督预训练 | 2026-09-21 |
| 9 | screenpipe screenpipe | 21.9k | 2.2k | Rust | 23 | YC (S26) | Open Computer History | Continuously record your company computer work, map your workflows, help you find work worth automating, and power your agents' context | 2026-10-11 |
| 10 | serve jina-ai | 21.9k | 2.2k | Python | 0 | ☁️ Build multimodal AI applications with cloud-native stack☁️ 使用云原生堆栈构建多模式人工智能应用程序 | 2025-03-24 |
| 11 | Qwen3-VL QwenLM | 20.1k | 1.9k | Jupyter Notebook | 390 | Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.Qwen3-VL是阿里云Qwen团队开发的多模态大语言模型系列。 | 2026-01-30 |
| 12 | Speech NVIDIA-NeMo | 18.6k | 3.6k | Python | 136 | A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)一个可扩展的生成式人工智能框架,专为从事大型语言模型、多模式和语音人工智能(自动语音识别和文本转语音)工作的研究人员和开发人员而构建 | 2026-10-10 |
| 13 | Awesome-Multimodal-Large-Language-Models BradyFU | 18.1k | 1.1k | N/A | 48 | :sparkles::sparkles:Latest Advances on Multimodal Large Language Models:sparkles::sparkles:多模态大语言模型的最新进展 | 2026-10-01 |
| 14 | Janus deepseek-ai | 17.8k | 2.2k | Python | 160 | Janus-Series: Unified Multimodal Understanding and Generation ModelsJanus 系列:统一多模态理解和生成模型 | 2025-02-01 |
| 15 | pipecat pipecat-ai | 16.3k | 2.9k | Python | 113 | Open Source framework for voice agents, multimodal apps, and realtime AI. Maintained by Daily and the community.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-10-11 |
| 16 | ms-swift modelscope | 15.8k | 1.7k | Python | 456 | Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.8, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM5.3, Gemma4, Llava, Phi4, ...) (AAAI 2025). | 2026-10-11 |
| 17 | all-in-rag datawhalechina | 11.9k | 5.9k | Python | 19 | 🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/ | 2026-09-30 |
| 18 | lancedb lancedb | 11.6k | 1.1k | Rust | 510 | Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.适用于多模式 AI 的开发人员友好型 OSS 嵌入式检索库。搜索更多;少管理。 | 2026-10-10 |
| 19 | rerun rerun-io | 11.6k | 863 | Rust | 1221 | Visualize, query, and stream to train on multimodal robotics data.可视化、查询和流式传输以训练多模式机器人数据。 | 2026-10-10 |
| 20 | X-AnyLabeling CVHub520 | 10.6k | 1.2k | Python | 5 | X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.X-AnyLabeling:一款轻量级、高效且统一的跨平台桌面应用程序,用于注释文本、图像、视频和多模式数据,将多功能内置工具与最先进的人工智能模型和灵活的多格式导出相结合。 | 2026-10-06 |
| 21 | runanywhere-sdks RunanywhereAI | 10.3k | 388 | C++ | 83 | Production ready toolkit to run AI locally用于本地运行 AI 的生产就绪工具包 | 2026-10-09 |
| 22 | self-operating-computer OthersideAI | 10.3k | 1.4k | Python | 83 | A framework to enable a multimodal model to operate a computer.使多模式模型能够操作计算机的框架。 | 2025-09-19 |
| 23 | PixelRAG StarTrail-org | 10.2k | 897 | Python | 8 | https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-10-08 |
| 24 | pyod yzhao062 | 10.0k | 1.5k | Python | 200 | A Python library for anomaly detection across tabular, time series, graph, text, image, and audio data. 60+ detectors, benchmark-backed ADEngine orchestration, and an agentic workflow for AI agents.用于跨表格、时间序列、图形、文本、图像和音频数据进行异常检测的 Python 库。 60 多个检测器、基准支持的 ADEngine 编排以及 AI 代理的代理工作流程。 | 2026-10-04 |
| 25 | gorse gorse-io | 9.9k | 918 | Go | 104 | AI powered open source recommender system engine supports classical/LLM rankers and multimodal content via embedding人工智能驱动的开源推荐系统引擎通过嵌入支持经典/LLM 排名和多模式内容 | 2026-10-10 |
| 26 | AIMLInterviews alirezadir | 9.8k | 1.7k | Jupyter Notebook | 5 | A Comprehensive Guide for Generative AI, Agentic AI, Multimodal AI, Physical AI Interviews. Proven Formula for multiple FAANG offers and more | 2026-10-08 |
| 27 | seatunnel apache | 9.7k | 2.4k | Java | 483 | SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.SeaTunnel是一个多模态、高性能、分布式、海量数据集成工具。 | 2026-10-09 |
| 28 | inference xorbitsai | 9.6k | 875 | Python | 5 | Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.通过更改一行代码即可将 GPT 替换为任何 LLM。 Xinference 可让您在云、本地或笔记本电脑上运行开源、语音和多模式模型 — 所有这些都通过一个统一的、可用于生产的推理 API。 | 2026-10-11 |
| 29 | MobileAgent X-PLUG | 9.3k | 930 | Python | 192 | Mobile-Agent: The Powerful GUI Agent FamilyMobile-Agent:强大的 GUI 代理系列 | 2026-07-07 |
| 30 | deeplake activeloopai | 9.2k | 724 | C++ | 50 | Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.Deeplake 是代理的人工智能数据运行时。它为无服务器 postgres 提供多模式数据湖,从而实现可扩展的检索和训练。 | 2026-05-21 |
| 31 | BentoML bentoml | 8.9k | 1.0k | Python | 148 | The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!服务 AI 应用程序和模型的最简单方法 - 构建模型推理 API、作业队列、LLM 应用程序、多模型管道等等! | 2026-10-05 |
| 32 | mlx-audio Blaizzy | 8.0k | 748 | Python | 80 | A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.基于 Apple MLX 框架构建的文本转语音 (TTS)、语音转文本 (STT) 和语音转语音 (STS) 库,可在 Apple Silicon 上提供高效的语音分析。 | 2026-10-09 |
| 33 | mmagic open-mmlab | 7.5k | 1.1k | Jupyter Notebook | 62 | OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.OpenMMLab 多模式高级、生成和智能创建工具箱。解锁魔法🪄:生成式人工智能 (AIGC)、易于使用的 API、出色的模型动物园、扩散模型,用于文本到图像生成、图像/视频恢复/增强等。 | 2024-08-06 |
| 34 | lance lance-format | 7.1k | 876 | Rust | 819 | Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and PyTorch with more integrations coming..多模式 AI 的开放 Lakehouse 格式。只需 2 行代码即可从 Parquet 进行转换,以实现速度提高 100 倍的随机访问、向量索引和数据版本控制。与 Pandas、DuckDB、Polars、Pyarrow 和 PyTorch 兼容,即将推出更多集成。 | 2026-10-11 |
| 35 | vllm-omni vllm-project | 7.1k | 1.9k | Python | 881 | A framework for efficient model inference with omni-modality models全模态模型的高效模型推理框架 | 2026-10-11 |
| 36 | GLM-4 zai-org | 7.1k | 619 | Python | 40 | GLM-4 series: Open Multilingual Multimodal Chat LMs | 开源多语言多模态对话模型 | 2026-08-05 |
| 37 | awesome-multimodal-ml pliang279 | 6.9k | 895 | N/A | 7 | Reading list for research topics in multimodal machine learning多模态机器学习研究主题的阅读清单 | 2024-08-20 |
| 38 | AppAgent TencentQQGYLab | 6.9k | 757 | Python | 88 | AppAgent: Multimodal Agents as Smartphone Users, an LLM-based multimodal agent framework designed to operate smartphone apps.AppAgent:作为智能手机用户的多模式代理,一个基于法学硕士的多模式代理框架,旨在操作智能手机应用程序。 | 2025-03-19 |
| 39 | jaaz 11cafe | 6.7k | 667 | TypeScript | 39 | The world's first open-source multimodal creative assistant This is a substitute for Canva and Manus that prioritizes privacy and is usable locally.全球首款开源多模态创意助手 这是 Canva 和 Manus 的替代品,优先考虑隐私且可在本地使用。 | 2026-03-02 |
| 40 | podcastfy souzatharsis | 6.6k | 765 | Python | 87 | An Open Source Python alternative to NotebookLM's podcast feature: Transforming Multimodal Content into Captivating Multilingual Audio Conversations with GenAINotebookLM 播客功能的开源 Python 替代方案:使用 GenAI 将多模式内容转换为迷人的多语言音频对话 | 2026-05-04 |
| 41 | genkit genkit-ai | 6.5k | 859 | TypeScript | 548 | Open-source framework for building agentic apps in JavaScript, Go, Dart, and Python, built and used in production by Google用于在 JavaScript、Go、Dart 和 Python 中构建代理应用程序的开源框架,由 Google 在生产中构建和使用 | 2026-10-11 |
| 42 | courses SkalskiP | 6.5k | 596 | Python | 5 | This repository is a curated collection of links to various courses and resources about Artificial Intelligence (AI)该存储库是有关人工智能 (AI) 的各种课程和资源的链接的精选集合 | 2024-04-22 |
| 43 | ai-notes swyxio | 6.2k | 561 | HTML | 4 | notes for software engineers getting up to speed on new AI developments. Serves as datastore for https://latent.space writing, and product brainstorming, but has cleaned up canonical references under the /Resources folder.供软件工程师了解新的人工智能发展的笔记。用作 https://latent.space 写作和产品头脑风暴的数据存储,但已清理 /Resources 文件夹下的规范引用。 | 2026-02-16 |
| 44 | Bagel ByteDance-Seed | 6.2k | 543 | Python | 143 | Open-source unified multimodal model开源统一多式联运模型 | 2026-05-04 |
| 45 | VLM-R1 om-ai-lab | 6.0k | 384 | Python | 163 | Solve Visual Understanding with Reinforced VLMs使用增强型 VLM 解决视觉理解问题 | 2026-07-07 |
| 46 | pyspur PySpur-Dev | 5.8k | 430 | TypeScript | 30 | A visual playground for agentic workflows: Iterate over your agents 10x faster代理工作流程的可视化游乐场:代理迭代速度提高 10 倍 | 2026-06-29 |
| 47 | Daft Eventual-Inc | 5.8k | 568 | Rust | 288 | High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale适用于人工智能和多模式工作负载的高性能数据引擎。处理任何规模的图像、音频、视频和结构化数据 | 2026-10-08 |
| 48 | UltraRAG OpenBMB | 5.7k | 451 | Python | 7 | A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines用于构建复杂且创新的 RAG 管道的低代码 MCP 框架 | 2026-10-10 |
| 49 | mmf facebookresearch | 5.6k | 937 | Python | 115 | A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)Facebook AI Research (FAIR) 的视觉和语言多模态研究模块化框架 | 2026-10-06 |
| 50 | neuraltalk karpathy | 5.5k | 1.3k | Python | 26 | NeuralTalk is a Python+numpy project for learning Multimodal Recurrent Neural Networks that describe images with sentences.NeuralTalk 是一个 Python+numpy 项目,用于学习用句子描述图像的多模态循环神经网络。 | 2020-12-22 |
| 51 | DeepSeek-VL2 deepseek-ai | 5.4k | 1.8k | Python | 102 | DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal UnderstandingDeepSeek-VL2:用于高级多模态理解的专家混合视觉语言模型 | 2025-02-26 |
| 52 | xtuner InternLM | 5.2k | 455 | Python | 245 | A Next-Generation Training Engine Built for Ultra-Large MoE Models专为超大型 MoE 模型打造的下一代训练引擎 | 2026-10-10 |
| 53 | align-anything PKU-Alignment | 4.7k | 503 | Python | 28 | Align Anything: Training All-modality Model with Feedback对齐一切:通过反馈训练全模态模型 | 2025-11-27 |
| 54 | tree-of-thoughts kyegomez | 4.6k | 375 | Python | 0 | Plug in and Play Implementation of Tree of Thoughts: Deliberate Problem Solving with Large Language Models that Elevates Model Reasoning by atleast 70% 即插即用实现思想之树:使用大型语言模型深思熟虑地解决问题,将模型推理能力提升至少 70% | 2026-09-29 |
| 55 | ultravox fixie-ai | 4.6k | 385 | Python | 54 | A fast multimodal LLM for real-time voice用于实时语音的快速多模式法学硕士 | 2025-12-12 |
| 56 | Awesome-AIGC-Tutorials luban-agi | 4.6k | 295 | N/A | 5 | Curated tutorials and resources for Large Language Models, AI Painting, and more. 针对大型语言模型、AI 绘画等的精选教程和资源。 | 2024-03-31 |
| 57 | img2dataset rom1504 | 4.5k | 379 | Python | 125 | Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.轻松将大量图像 URL 转换为图像数据集。可以在一台机器上 20 小时内下载、调整大小和打包 100M 网址。 | 2025-10-19 |
| 58 | ruby_llm crmne | 4.5k | 512 | Ruby | 3 | The Ruby-native AI framework. Chats, agents, tools, images, audio, and video through one consistent API, in plain Ruby or Rails. | 2026-10-08 |
| 59 | lmms-eval EvolvingLMMs-Lab | 4.4k | 672 | Python | 26 | One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks跨文本、图像、视频和音频任务的一站式多模态评估工具包 | 2026-10-10 |
| 60 | modlens liustack | 4.2k | 129 | TypeScript | 2 | The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-10-04 |
| 61 | MOSS-TTS OpenMOSS | 4.2k | 379 | Python | 17 | An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS用于长篇语音、对话合成、语音设计、音效和实时流式 TTS 的开源模型系列 | 2026-09-06 |
| 62 | VisualGLM-6B zai-org | 4.2k | 419 | Python | 269 | Chinese and English multimodal conversational language model | 多模态中英双语对话语言模型 | 2024-08-23 |
| 63 | Fengshenbang-LM IDEA-CCNL | 4.1k | 371 | Python | 104 | Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。 | 2026-06-08 |
| 64 | open_flamingo mlfoundations | 4.1k | 320 | Python | 45 | An open-source framework for training large multimodal models.用于训练大型多模式模型的开源框架。 | 2024-08-31 |
| 65 | OmniGen2 VectorSpaceLab | 4.1k | 36 | Jupyter Notebook | 99 | OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871OmniGen2:对高级多模式生成的探索。 https://arxiv.org/abs/2506.18871 | 2026-03-20 |
| 66 | Qwen2.5-Omni QwenLM | 4.1k | 331 | Jupyter Notebook | 215 | Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-06-12 |
| 67 | mm-cot amazon-science | 4.0k | 330 | Python | 44 | Official implementation for "Multimodal Chain-of-Thought Reasoning in Language Models" (stay tuned and more will be updated)《语言模型中的多模态思维链推理》正式实现(敬请期待,更多内容将会更新) | 2024-06-12 |
| 68 | VILA NVlabs | 3.9k | 335 | Python | 67 | VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.VILA 是一系列最先进的视觉语言模型 (VLM),适用于跨边缘、数据中心和云的各种多模式 AI 任务。 | 2026-03-12 |
| 69 | mmpretrain open-mmlab | 3.9k | 1.1k | Python | 205 | OpenMMLab Pre-training Toolbox and BenchmarkOpenMMLab 预训练工具箱和基准测试 | 2024-11-01 |
| 70 | SimpleMem aiming-lab | 3.8k | 404 | Python | 3 | [ICML'26] SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal | 2026-07-24 |
| 71 | discoart jina-ai | 3.8k | 241 | Python | 28 | 🪩 Create Disco Diffusion artworks in one line🪩 用一行创建 Disco Diffusion 艺术品 | 2023-05-16 |
| 72 | morphik-core morphik-org | 3.7k | 328 | Python | 14 | Open-source multimodal retrieval engine (Morphik Core) | 2026-10-05 |
| 73 | Awesome-LLM-Reasoning atfortes | 3.7k | 215 | N/A | 5 | From Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-04-20 |
| 74 | NExT-GPT NExT-GPT | 3.6k | 359 | Python | 81 | Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language ModelError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-05-13 |
| 75 | mini-omni gpt-omni | 3.6k | 311 | Python | 37 | open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities. Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2024-11-05 |
| 76 | awesome-embodied-vla-va-vln jonyzhang2023 | 3.6k | 173 | N/A | 1 | A curated list of state-of-the-art research in embodied AI, focusing on vision-language-action (VLA) models, vision-language navigation (VLN), and related multimodal learning approaches. 嵌入式人工智能最先进研究的精选列表,重点关注视觉-语言-动作 (VLA) 模型、视觉-语言导航 (VLN) 和相关的多模态学习方法。 | 2026-09-15 |
| 77 | py-xiaozhi huangjunsen0406 | 3.5k | 745 | Python | 6 | Open-source AI assistant ecosystem with MCP integrations, multimodal workflows, IoT support, and cross-platform voice interaction.开源人工智能助手生态系统,具有 MCP 集成、多模式工作流程、物联网支持和跨平台语音交互。 | 2026-09-11 |
| 78 | mteb embeddings-benchmark | 3.5k | 719 | Python | 281 | MTEB: State-of-the-art evaluation of embeddings across languages and modalitiesMTEB:跨语言和模式嵌入的最先进评估 | 2026-10-11 |
| 79 | claude-code-local nicedreamzapp | 3.4k | 633 | Python | 3 | Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl. Muse-Glimmer 30B (now multimodal — reads images, abliterated), Gemma 4 31B, Qwen 3.5 122B (65 tok/s), DeepSeek V4 Flash (1M ctx). Private, offline, airgap-ready. Built for NDA / legal / healthcare workflows.使用 Apple Silicon 上的本地 AI 在设备上 100% 运行 Claude Code。 MLX 原生 Anthropic-API 服务器。包括 6 名战斗机Muse-Glimmer 30B(现在是多模式 — 读取图像,已删除)、Gemma 4 31B、Qwen 3.5 122B (65 tok/s)、DeepSeek V4 Flash (1M ctx)。私密、离线、气隙就绪。专为 NDA/法律/医疗保健工作流程而构建。 | 2026-10-09 |
| 80 | HunyuanImage-3.0 Tencent-Hunyuan | 3.3k | 191 | Python | 45 | HunyuanImage-3.0: A Powerful Native Multimodal Model for Image GenerationHunyuanImage-3.0:强大的原生图像生成多模态模型 | 2026-06-23 |
| 81 | tribev2 facebookresearch | 3.3k | 707 | Jupyter Notebook | 23 | This repository contains the code to train and evaluate TRIBE v2, a multimodal model for brain response prediction该存储库包含训练和评估 TRIBE v2 的代码,TRIBE v2 是一种用于大脑反应预测的多模式模型 | 2026-06-23 |
| 82 | LakeSoul lakesoul-io | 3.3k | 429 | Rust | 18 | LakeSoul is an end-to-end, realtime cloud-native Lakehouse framework for fast data ingestion, concurrent updates, incremental analytics, multimodal data processing and vector search — powering next-generation BI and AI workloads.LakeSoul 是一种端到端、实时云原生 Lakehouse 框架,用于快速数据摄取、并发更新、增量分析、多模式数据处理和矢量搜索,为下一代 BI 和 AI 工作负载提供动力。 | 2026-10-11 |
| 83 | vortex vortex-data | 3.3k | 239 | Rust | 238 | An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.可扩展、最先进的列式压缩框架,以及最快的 FOSS 列式文件格式。以前在 @spiraldb,现在是 LFAI&Data(Linux 基金会的一部分)的孵化阶段项目。 | 2026-10-10 |
| 84 | InternGPT OpenGVLab | 3.2k | 231 | Python | 19 | InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models.现在它支持DragGAN、ChatGPT、ImageBind、多模式聊天(如GPT-4、SAM)、交互式图像编辑等。请在igpt.opengvlab.com上尝试(支持DragGAN、ChatGPT、ImageBind、SAM的在线演示系统) | 2024-08-20 |
| 85 | OSWorld xlang-ai | 3.2k | 540 | Python | 158 | [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments[NeurIPS 2024] OSWorld:真实计算机环境中开放式任务的多模式代理基准测试 | 2026-09-14 |
| 86 | ai TanStack | 3.2k | 361 | TypeScript | 24 | 🤖 Type-safe, provider-agnostic TypeScript AI SDK for streaming chat, tool calling, agents, and multimodal apps across OpenAI, Anthropic, Gemini, React, Vue, Svelte, and Solid.🤖 类型安全、与提供商无关的 TypeScript AI SDK,用于跨 OpenAI、Anthropic、Gemini、React、Vue、Svelte 和 Solid 的流式聊天、工具调用、代理和多模式应用程序。 | 2026-10-10 |
| 87 | Skywork-R1V SkyworkAI | 3.2k | 284 | Python | 35 | Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in vision-language reasoning.Skywork-R1V是Skywork AI开发的先进多模态AI模型系列,专注于视觉语言推理。 | 2026-07-29 |
| 88 | torchscale microsoft | 3.1k | 224 | Python | 30 | Foundation Architecture for (M)LLMs(M)LLM 的基础架构 | 2024-04-11 |
| 89 | Qwen-MM-Plugins QwenLM | 3.1k | 207 | Python | 8 | Make any agent harness multimodal-native.使任何代理利用多模式原生。 | 2026-10-10 |
| 90 | docarray docarray | 3.1k | 241 | Python | 68 | Represent, send, store and search multimodal data表示、发送、存储和搜索多模式数据 | 2026-03-27 |
| 91 | helm stanford-crfm | 2.9k | 421 | Python | 67 | Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.语言模型的整体评估 (HELM) 是由斯坦福大学基础模型研究中心 (CRFM) 创建的开源 Python 框架,用于对基础模型进行整体、可重复和透明的评估,包括大语言模型 (LLM) 和多模态模型。 | 2026-09-01 |
| 92 | InternLM-XComposer InternLM | 2.9k | 177 | Python | 139 | InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio InteractionsInternLM-XComposer2.5-OmniLive:用于长期流媒体视频和音频交互的综合多模式系统 | 2025-05-26 |
| 93 | Awesome-AI4Med FreedomIntelligence | 2.9k | 497 | N/A | 0 | A curated list of medical LLMs, multimodal systems, datasets, benchmarks, and more. 🏥医学法学硕士、多模式系统、数据集、基准等的精选列表。 🏥 | 2026-08-04 |
| 94 | datachain datachain-ai | 2.8k | 166 | Python | 81 | The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure非结构化数据的上下文层:S3、GCS、Azure 上的类型化、版本化数据集 | 2026-10-11 |
| 95 | clip-retrieval rom1504 | 2.8k | 236 | Jupyter Notebook | 79 | Easily compute clip embeddings and build a clip retrieval system with them轻松计算剪辑嵌入并用它们构建剪辑检索系统 | 2026-03-28 |
| 96 | autodistill autodistill | 2.8k | 226 | Python | 40 | Images to inference with no labeling (use foundation models to train supervised models).无需标记即可进行推理的图像(使用基础模型来训练监督模型)。 | 2026-09-29 |
| 97 | MUNIT NVlabs | 2.7k | 487 | Python | 63 | Multimodal Unsupervised Image-to-Image Translation多模态无监督图像到图像翻译 | 2022-09-20 |
| 98 | maestro roboflow | 2.7k | 225 | Python | 17 | streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL简化多模式模型的微调过程:PaliGemma 2、Florence-2 和 Qwen2.5-VL | 2026-10-05 |
| 99 | OmAgent om-ai-lab | 2.7k | 291 | Python | 11 | [EMNLP-2024] Build multimodal language agents for fast prototype and production[EMNLP-2024] 构建多模式语言代理以实现快速原型和生产 | 2025-03-19 |
| 100 | generative-ai genieincodebottle | 2.6k | 642 | Jupyter Notebook | 0 | Comprehensive resources on Generative AI, including a detailed roadmap, projects, use cases, interview preparation, and coding preparation.关于生成人工智能的综合资源,包括详细的路线图、项目、用例、面试准备和编码准备。 | 2026-10-08 |
No repositories match your search
没有匹配的仓库