Back to Rankings返回排行榜
Top 100 · RLHF / alignment前 100 · RLHF / 对齐
100 repositories sorted by rlhf / alignment 按 RLHF / 对齐 排序,共 100 个仓库
| # | Repository仓库 | Stars | Forks | Language语言 | Issues | Description描述 | Last Commit最后提交 |
|---|---|---|---|---|---|---|---|
| 1 | LlamaFactory hiyouga | 75.4k | 9.2k | Python | 1004 | Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)100多个LLM和VLM的统一高效微调(ACL 2024) | 2026-10-10 |
| 2 | Open-Assistant LAION-AI | 37.4k | 3.3k | Python | 228 | OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.OpenAssistant 是一个基于聊天的助手,它可以理解任务,可以与第三方系统交互,并动态检索信息来执行此操作。 | 2024-08-17 |
| 3 | LLMSurvey RUCAIBox | 12.2k | 930 | Python | 26 | The official GitHub page for the survey paper "A Survey of Large Language Models".Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-03-11 |
| 4 | OpenRLHF OpenRLHF | 10.1k | 1.0k | Python | 309 | An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)基于 Ray 的易于使用、可扩展且高性能的 Agentic RL 框架(PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL) | 2026-10-05 |
| 5 | PaLM-rlhf-pytorch lucidrains | 7.9k | 670 | Python | 17 | Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the PaLM architecture. Basically ChatGPT but with PaLMImplementation of RLHF (Reinforcement Learning with Human Feedback) on top of the PaLM architecture.基本上是 ChatGPT,但使用 PaLM | 2026-09-20 |
| 6 | InternLM InternLM | 7.3k | 511 | Python | 8 | Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-10-30 |
| 7 | Chinese-LLaMA-Alpaca-2 ymcui | 7.1k | 556 | Python | 1 | 中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models) | 2026-04-19 |
| 8 | MedicalGPT shibing624 | 5.9k | 800 | Python | 6 | MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型,实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。 | 2026-09-15 |
| 9 | OpenClaw-RL Gen-Verse | 5.7k | 613 | Python | 51 | OpenClaw-RL: Train any agent simply by talkingError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-05-23 |
| 10 | alignment-handbook huggingface | 5.7k | 496 | Python | 92 | Robust recipes to align language models with human and AI preferencesError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-09-23 |
| 11 | reasoning-from-scratch rasbt | 5.4k | 855 | Jupyter Notebook | 2 | Implement a reasoning LLM in PyTorch from scratch, step by step从头开始,一步步在 PyTorch 中实现推理 LLM | 2026-10-02 |
| 12 | transformerlab-app transformerlab | 5.2k | 551 | Python | 13 | The open source research environment for AI researchers to seamlessly train, evaluate, and scale models from local hardware to GPU clusters.供 AI 研究人员无缝训练、评估和扩展从本地硬件到 GPU 集群的模型的开源研究环境。 | 2026-10-03 |
| 13 | Kiln Kiln-AI | 5.2k | 392 | Python | 23 | Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-10-09 |
| 14 | argilla argilla-io | 5.1k | 506 | Python | 9 | Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasetsError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-10-06 |
| 15 | trlx CarperAI | 4.8k | 486 | Python | 86 | A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)通过人类反馈(RLHF)进行强化学习的语言模型分布式训练的存储库 | 2024-01-08 |
| 16 | align-anything PKU-Alignment | 4.7k | 503 | Python | 28 | Align Anything: Training All-modality Model with Feedback对齐一切:通过反馈训练全模态模型 | 2025-11-27 |
| 17 | hands-on-modern-rl walkinglabs | 4.6k | 334 | Python | 11 | 🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems. 🚀 一个开源的实践课程,弥合了从基本 RL 概念到 LLM 对齐、RLVR 和高级 Agentic 系统的差距。 | 2026-10-01 |
| 18 | awesome-RLHF opendilab | 4.4k | 263 | N/A | 0 | A curated list of reinforcement learning with human feedback resources (continually updated)带有人类反馈资源的强化学习精选列表(持续更新) | 2026-05-20 |
| 19 | ChatGLM-Efficient-Tuning hiyouga | 3.7k | 464 | Python | 0 | Fine-tuning ChatGLM-6B with PEFT | 基于 PEFT 的高效 ChatGLM 微调 | 2023-10-12 |
| 20 | docta Docta-ai | 3.5k | 256 | Python | 1 | A Doctor for your data您的数据医生 | 2026-06-16 |
| 21 | ROLL alibaba | 3.4k | 315 | Python | 102 | An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language ModelsError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-10-11 |
| 22 | distilabel argilla-io | 3.4k | 260 | Python | 82 | Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.Distilabel 是一个合成数据和人工智能反馈框架,适用于需要基于经过验证的研究论文的快速、可靠和可扩展管道的工程师。 | 2026-10-06 |
| 23 | Cortex qibin0506 | 2.7k | 209 | Python | 0 | 从零构建大模型:从预训练到RLHF的完整实践 | 2026-08-21 |
| 24 | TorchLeet Exorust | 2.5k | 312 | Jupyter Notebook | 3 | LeetCode for PyTorch — 65 ML/AI interview problems from real interviews at Google, Meta, Anthropic. Jupyter notebooks, an auto-grader, and an MCP AI tutor.LeetCode for PyTorch — 65 个 ML/AI 面试问题,来自 Google、Meta、Anthropic 的真实面试。 Jupyter 笔记本、自动评分器和 MCP AI 导师。 | 2026-09-15 |
| 25 | rlhf-book natolambert | 2.4k | 282 | Python | 1 | Textbook on reinforcement learning from human feedbackError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-09-24 |
| 26 | transformers_tasks HarderThenHarder | 2.4k | 399 | Jupyter Notebook | 59 | ⭐️ NLP Algorithms with transformers lib. Supporting Text-Classification, Text-Generation, Information-Extraction, Text-Matching, RLHF, SFT etc.⭐️ 带有 Transformer lib 的 NLP 算法。 Supporting Text-Classification, Text-Generation, Information-Extraction, Text-Matching, RLHF, SFT etc. | 2023-09-29 |
| 27 | alpaca_eval tatsu-lab | 2.0k | 314 | Jupyter Notebook | 25 | An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.An automatic evaluator for instruction-following language models.经过人工验证、高质量、便宜且快速。 | 2025-08-09 |
| 28 | AgentsMeetRL thinkwee | 1.9k | 75 | HTML | 0 | Awesome List for Agentic RLAgentic RL 的精彩列表 | 2026-10-10 |
| 29 | hh-rlhf anthropics | 1.9k | 162 | N/A | 0 | Human preference data for "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback"Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-06-17 |
| 30 | ChatLM-mini-Chinese charent | 1.7k | 191 | Python | 10 | 中文对话0.2B小模型(ChatLM-Chinese-0.2B),开源所有数据集来源、数据清洗、tokenizer训练、模型预训练、SFT指令微调、RLHF优化等流程的全部代码。支持下游任务sft微调,给出三元组信息抽取微调示例。 | 2024-04-20 |
| 31 | ImageReward zai-org | 1.7k | 90 | Python | 58 | [NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation[NeurIPS 2023] ImageReward:学习和评估人类对文本到图像生成的偏好 | 2025-10-29 |
| 32 | safe-rlhf PKU-Alignment | 1.6k | 134 | Python | 17 | Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human FeedbackError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-11-24 |
| 33 | WebGLM THUDM | 1.6k | 130 | Python | 51 | WebGLM: An Efficient Web-enhanced Question Answering System (KDD 2023)WebGLM:高效的网络增强问答系统 (KDD 2023) | 2025-03-25 |
| 34 | RLHF-Reward-Modeling RLHFlow | 1.5k | 111 | Python | 21 | Recipes to train reward model for RLHF.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-04-24 |
| 35 | MOSS-RLHF OpenLMLab | 1.4k | 103 | Python | 39 | Secrets of RLHF in Large Language Models Part I: PPOError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2024-03-03 |
| 36 | xtreme1 xtreme1-io | 1.4k | 222 | TypeScript | 20 | Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.Xtreme1 是一款用于多模态数据训练的一体化数据标记和注释平台,支持 3D LiDAR 点云、图像和 LLM。 | 2026-09-29 |
| 37 | pyre-code whwangovo | 1.3k | 116 | Python | 4 | A self-hosted ML coding practice platform. 68 problems from ReLU to flow matching — attention, training, RLHF, diffusion, and more. Instant feedback in the browser.一个自托管的机器学习编码练习平台。 68 problems from ReLU to flow matching — attention, training, RLHF, diffusion, and more.浏览器中的即时反馈。 | 2026-05-12 |
| 38 | verl-omni verl-project | 1.2k | 238 | Python | 98 | Multimodal RL training framework for diffusion & omni models用于扩散和全向模型的多模态强化学习训练框架 | 2026-10-10 |
| 39 | labs-molt NVIDIA-NeMo | 1.2k | 113 | Python | 9 | A scalable, agentic-first, and HuggingFace-native RL framework for research (9k lines). | 2026-10-08 |
| 40 | AI-Compass tingaicompass | 976 | 128 | Python | 3 | “AI-Compass”将为社区指引在 AI 技术海洋中航行的方向,无论你是初学者还是进阶开发者,都能在这里找到通往 AI 各大方向的路径。旨在帮助开发者系统性地了解 AI 的核心概念、主流技术、前沿趋势,并通过实践掌握从理论到落地的全过程。 | 2026-09-24 |
| 41 | SimPO princeton-nlp | 961 | 78 | Python | 25 | [NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward[NeurIPS 2024] SimPO:具有无参考奖励的简单偏好优化 | 2025-02-16 |
| 42 | HALOs ContextualAI | 909 | 51 | Python | 7 | A library with extensible implementations of DPO, KTO, PPO, ORPO, and other human-aware loss functions (HALOs).具有 DPO、KTO、PPO、ORPO 和其他人类感知损失函数 (HALO) 的可扩展实现的库。 | 2025-09-30 |
| 43 | awesome-on-policy-distillation chrisliu298 | 877 | 38 | N/A | 0 | A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models用于大型语言模型的策略蒸馏 (OPD) 的论文、技术报告、框架和工具的精选集 | 2026-10-06 |
| 44 | OpenJudge agentscope-ai | 871 | 73 | Python | 14 | OpenJudge: A Unified Framework for Holistic Evaluation and Quality RewardsError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-09-11 |
| 45 | alpaca_farm tatsu-lab | 844 | 66 | Python | 6 | A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data. Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2024-07-01 |
| 46 | halo whitecircle | 820 | 39 | Python | 12 | Halo is an open-source framework built by White Circle for training large language and multimodal models | 2026-10-11 |
| 47 | reward-bench allenai | 743 | 103 | Python | 3 | RewardBench: the first evaluation tool for reward models.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-09-30 |
| 48 | AlignLLMHumanSurvey GaryYufei | 738 | 31 | N/A | 1 | Aligning Large Language Models with Human: A Survey使大型语言模型与人类保持一致:一项调查 | 2026-09-08 |
| 49 | Trinity-RFT agentscope-ai | 708 | 84 | Python | 41 | Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (LLM).Trinity-RFT 是一个通用、灵活且可扩展的框架,专为大语言模型 (LLM) 的强化微调 (RFT) 而设计。 | 2026-10-08 |
| 50 | LLM-Algorithm-Intern-Guide Junvate | 703 | 13 | N/A | 0 | 🚀 2026届大模型算法岗实习面经 | 包含 DeepSeek/Qwen 技术报告解析、手撕 PPO/RoPE/Transformer、RLHF 核心与八股文 | 持续更新中... | 2026-03-28 |
| 51 | oat sail-sg | 673 | 63 | Python | 6 | 🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-01-29 |
| 52 | Cornucopia-LLaMA-Fin-Chinese jerry1993-tech | 655 | 66 | Python | 17 | 聚宝盆(Cornucopia): 中文金融系列开源可商用大模型,并提供一套高效轻量化的垂直领域LLM训练框架(Pretraining、SFT、RLHF、Quantize等) | 2023-06-30 |
| 53 | Open-AgentRL Gen-Verse | 651 | 60 | Python | 10 | RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic ScenariosRLAnything (ICML 2026) 和 AutoTool (ICML 2026)、DemyAgent:适用于法学硕士和代理场景的开源强化学习 | 2026-06-12 |
| 54 | Relax redai-studio | 641 | 177 | Python | 57 | An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale用于大规模全模态后训练的异步强化学习引擎 | 2026-10-10 |
| 55 | LLamaTuner jianzhnie | 622 | 62 | Python | 18 | Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署. | 2025-01-24 |
| 56 | SPPO uclaml | 588 | 48 | Python | 14 | The official implementation of Self-Play Preference Optimization (SPPO)Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-01-23 |
| 57 | ARIS-in-AI-Offer wanshuiyin | 586 | 22 | Python | 3 | Bilingual (中文+EN) ML / LLM / diffusion / agent interview cheat sheets for AI 秋招 — generated by ARIS /interview-cheatsheet, rendered by /render-html into single-file HTML, reads anywhere — plus a CV→DBLP-fact-checked academic homepage generator and hand-authored long-form blogs 🌱双语 (中文+EN) ML / LLM / 扩散 / 代理 AI 秋招面试备忘单 — 由 ARIS /interview-cheatsheet 生成,由 /render-html 渲染为单文件 HTML,可在任何地方阅读 — 加上经过 CV→DBLP 事实检查的学术主页生成器和手工撰写的长篇博客 🌱 | 2026-10-06 |
| 58 | Awesome-LLM-On-Policy-Distillation nick7nlp | 575 | 16 | Python | 2 | A curated collection of papers and resources on On-Policy Distillation for Large Language Models.有关大型语言模型的策略蒸馏的精选论文和资源集。 | 2026-10-06 |
| 59 | TextRL voidful | 565 | 61 | Python | 3 | Implementation of ChatGPT RLHF (Reinforcement Learning with Human Feedback) on any generation model in huggingface's transformer (blommz-176B/bloom/gpt/bart/T5/MetaICL)Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-04-23 |
| 60 | Online-RLHF RLHFlow | 545 | 48 | Python | 12 | A recipe for online RLHF and online iterative DPO.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2024-12-28 |
| 61 | dLLM-RL Gen-Verse | 523 | 45 | Python | 24 | [ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA TraDo series.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-01-28 |
| 62 | step_into_llm candle-org | 482 | 127 | Jupyter Notebook | 27 | MindSpore online courses: Step into LLMError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-12-22 |
| 63 | LaMDA-rlhf-pytorch conceptofmind | 467 | 73 | Python | 6 | Open-source pre-training implementation of Google's LaMDA in PyTorch. Adding RLHF similar to ChatGPT.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2024-02-24 |
| 64 | FineEnvs adithya-s-k | 463 | 61 | Python | 3 | FineEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs | 2026-10-08 |
| 65 | LLM-RLHF-Tuning Joyce94 | 452 | 24 | Python | 3 | LLM Tuning with PEFT (SFT+RM+PPO+DPO with LoRA) Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2023-10-11 |
| 66 | ABC-GRPO chi2liu | 442 | 42 | Python | 0 | Code For All-Quadrant Bounded Clipping GRPO. arxiv.org/pdf/2601.03895全象限限幅裁剪 GRPO 的代码。 arxiv.org/pdf/2601.03895 | 2026-09-08 |
| 67 | mlx-lm-lora Goekdeniz-Guelmez | 429 | 60 | Python | 0 | Train Large Language Models on MLX.在 MLX 上训练大型语言模型。 | 2026-10-05 |
| 68 | VisionReward zai-org | 427 | 15 | Python | 19 | [AAAI 2026] VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation[AAAI 2026] VisionReward:用于图像和视频生成的细粒度多维人类偏好学习 | 2025-03-26 |
| 69 | pykoi CambioML | 410 | 44 | Jupyter Notebook | 2 | pykoi: Active learning in one unified interfacepykoi:在一个统一的界面中进行主动学习 | 2025-09-24 |
| 70 | JarvisEvo LYL1015 | 408 | 9 | Python | 1 | [CVPR' 2026] JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator OptimizationError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2026-02-22 |
| 71 | LLaVA-RLHF llava-rlhf | 399 | 30 | Python | 4 | Aligning LMMs with Factually Augmented RLHF将 LMM 与事实增强的 RLHF 结合起来 | 2023-11-01 |
| 72 | quick-start-guide-to-llms sinanuozdemir | 397 | 213 | Jupyter Notebook | 1 | The Official Repo for "Quick Start Guide to Large Language Models"“大型语言模型快速入门指南”的官方存储库 | 2025-10-07 |
| 73 | awesome-llm-human-preference-datasets glgh | 394 | 19 | N/A | 0 | A curated list of Human Preference Datasets for LLM fine-tuning, RLHF, and eval.用于 LLM 微调、RLHF 和评估的人类偏好数据集的精选列表。 | 2023-10-04 |
| 74 | mLoRA TUDB-Labs | 385 | 70 | Python | 12 | An Efficient "Factory" to Build Multiple LoRA AdaptersError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-02-13 |
| 75 | AdaRubrics alphadl | 365 | 37 | Python | 1 | AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent TrajectoriesAdaRubric:智能体轨迹的自适应动态评估器 | 2026-08-24 |
| 76 | Stable-Alignment agi-templar | 356 | 18 | Python | 4 | Multi-agent Social Simulation + Efficient, Effective, and Stable alternative of RLHF. Code for the paper "Training Socially Aligned Language Models in Simulated Human Society".Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2023-06-18 |
| 77 | MedQA-ChatGLM WangRongsheng | 340 | 50 | Python | 3 | 🛰️ 基于真实医疗对话数据在ChatGLM上进行LoRA、P-Tuning V2、Freeze、RLHF等微调,我们的眼光不止于医疗问答 | 2023-09-02 |
| 78 | ReaLHF openpsi-project | 336 | 22 | Python | 0 | Super-Efficient RLHF Training of LLMs with Parameter Reallocation通过参数重新分配对 LLM 进行超高效 RLHF 训练 | 2025-04-24 |
| 79 | VADER mihirp1998 | 318 | 15 | Python | 11 | Video Diffusion Alignment via Reward Gradients. We improve a variety of video diffusion models such as VideoCrafter, OpenSora, ModelScope and StableVideoDiffusion by finetuning them using various reward models such as HPS, PickScore, VideoMAE, VJEPA, YOLO, Aesthetics etc. 通过奖励梯度进行视频扩散对齐。我们通过使用 HPS、PickScore、VideoMAE、VJEPA、YOLO、Aesthetics 等各种奖励模型进行微调,改进了各种视频扩散模型,例如 VideoCrafter、OpenSora、ModelScope 和 StableVideoDiffusion。 | 2025-03-12 |
| 80 | RLHF-V RLHF-V | 310 | 8 | Python | 2 | [CVPR'24] RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback[CVPR'24] RLHF-V:通过细粒度矫正人类反馈的行为调整,迈向值得信赖的 MLLM | 2024-09-11 |
| 81 | RLLoggingBoard HarderThenHarder | 296 | 9 | Python | 0 | A visuailzation tool to make deep understaning and easier debugging for RLHF training.一种可视化工具,可深入理解 RLHF 训练并更轻松地进行调试。 | 2025-02-20 |
| 82 | RLHF sunzeyeah | 284 | 35 | Python | 3 | Implementation of Chinese ChatGPTError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2023-11-20 |
| 83 | FineGrainedRLHF allenai | 283 | 24 | Python | 2 | 2025-01-06 | |
| 84 | unsloth-buddy TYH-labs | 281 | 15 | Python | 0 | Zero-friction LLM fine-tuning skill for Claude Code, Gemini CLI & any ACP agent. Unsloth on NVIDIA · TRL+MPS/MLX on Apple Silicon. Automates env setup, LoRA training (SFT, DPO, GRPO, vision), post-hoc GRPO log diagnostics, evaluation, and export end-to-end. Part of the Gaslamp AI platform.Zero-friction LLM fine-tuning skill for Claude Code, Gemini CLI & any ACP agent. Unsloth on NVIDIA · TRL+MPS/MLX on Apple Silicon. Automates env setup, LoRA training (SFT, DPO, GRPO, vision), post-hoc GRPO log diagnostics, evaluation, and export end-to-end. Gaslamp AI 平台的一部分。 | 2026-06-28 |
| 85 | MLX-LoRA-Studio Goekdeniz-Guelmez | 277 | 28 | Swift | 0 | A native Mac App for LLM fine-tuning on Apple Silicon — fully on-device, fully open source.用于在 Apple Silicon 上进行 LLM 微调的本机 Mac 应用程序 — 完全在设备上、完全开源。 | 2026-08-26 |
| 86 | Open-R1 jianzhnie | 273 | 55 | Python | 0 | The open source implementation of DeepSeek-R1. 开源复现 DeepSeek-R1 | 2025-03-10 |
| 87 | RLHF_in_notebooks ash80 | 254 | 33 | Jupyter Notebook | 0 | RLHF (Supervised fine-tuning, reward model, and PPO) step-by-step in 3 Jupyter notebooksError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-06-20 |
| 88 | RLHF-Label-Tool SupritYoung | 254 | 21 | Python | 2 | 用于大模型 RLHF 进行人工数据标注排序的工具。A tool for manual response data annotation sorting in RLHF stage.用于大模型 RLHF 进行人工数据标注排序的工具。 A tool for manual response data annotation sorting in RLHF stage. | 2023-08-01 |
| 89 | learn-MedicalGPT bcefghj | 244 | 14 | TypeScript | 1 | 🏥 从零基础到面试通关:20节课彻底搞懂MedicalGPT医疗大模型训练全流程 | PT/SFT/LoRA/RLHF/DPO/GRPO | 100+面试高频考点 | 2026-04-01 |
| 90 | llama-trl jasonvanf | 238 | 24 | Python | 7 | LLaMA-TRL: Fine-tuning LLaMA with PPO and LoRAError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-08-17 |
| 91 | LLaVA-MoD shufangxun | 228 | 17 | Python | 3 | [ICLR 2025] LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge DistillationError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2025-03-31 |
| 92 | RLHF HumanSignal | 227 | 43 | Jupyter Notebook | 3 | Collection of links, tutorials and best practices of how to collect the data and build end-to-end RLHF system to finetune Generative AI models关于如何收集数据和构建端到端 RLHF 系统以微调生成式 AI 模型的链接、教程和最佳实践的集合 | 2023-07-24 |
| 93 | chain-of-hindsight haoliuhl | 227 | 16 | Python | 3 | Simple next-token-prediction for RLHFError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2023-09-30 |
| 94 | minChatGPT ethanyanjiali | 226 | 37 | Python | 3 | A minimum example of aligning language models with RLHF similar to ChatGPTError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2023-09-26 |
| 95 | minimind-deep-dive Enping-Hu | 224 | 14 | Python | 0 | 从 MiniMind 源码读起,再延伸到现代大模型技术体系的中文学习笔记。主线逐行精读预训练 / SFT / DPO / PPO / GRPO 与训练机制;附录 17 篇进阶卷覆盖量化、投机解码、RLHF 全景、模型代际史等 MiniMind 没涉及、但进阶绕不开的主题。 | 2026-07-21 |
| 96 | rl-handbook lubludrova | 221 | 13 | MDX | 1 | A comprehensive guide to Reinforcement Learning | 2026-10-02 |
| 97 | Vicuna-LoRA-RLHF-PyTorch jackaduma | 220 | 18 | Python | 15 | A full pipeline to finetune Vicuna LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the Vicuna architecture. Basically ChatGPT but with Vicuna在消费类硬件上使用 LoRA 和 RLHF 微调 Vicuna LLM 的完整流程。在 Vicuna 架构之上实现 RLHF(带有人类反馈的强化学习)。 Basically ChatGPT but with Vicuna | 2024-05-20 |
| 98 | awesome-RLAIF mengdi-li | 210 | 6 | N/A | 0 | A continually updated list of literature on Reinforcement Learning from AI Feedback (RLAIF) 不断更新的关于人工智能反馈强化学习 (RLAIF) 的文献列表 | 2026-09-17 |
| 99 | nanoRLHF hyunwoongko | 204 | 19 | Python | 0 | nanoRLHF: from-scratch journey into how LLMs and RLHF really work.nanoRLHF:从头开始了解法学硕士和 RLHF 的真正运作方式。 | 2026-10-07 |
| 100 | VL-RLHF TideDra | 204 | 8 | Python | 11 | A RLHF Infrastructure for Vision-Language ModelsError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. | 2024-11-15 |
No repositories match your search
没有匹配的仓库