Github Ranking /
2026-10-11
Back to Rankings返回排行榜

⚖️ Top 100 · RLHF / alignment前 100 · RLHF / 对齐

100 repositories sorted by rlhf / alignment 按 RLHF / 对齐 排序,共 100 个仓库

⌕
📦 100 repos个仓库 🕐 2026-10-11
# Repository仓库 Stars Forks Language语言 Issues Description描述 Last Commit最后提交
1 LlamaFactory hiyouga 75.4k 9.2k Python 1004 Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)100多个LLM和VLM的统一高效微调(ACL 2024) 2026-10-10
2 Open-Assistant LAION-AI 37.4k 3.3k Python 228 OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.OpenAssistant 是一个基于聊天的助手,它可以理解任务,可以与第三方系统交互,并动态检索信息来执行此操作。 2024-08-17
3 LLMSurvey RUCAIBox 12.2k 930 Python 26 The official GitHub page for the survey paper "A Survey of Large Language Models".Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2025-03-11
4 OpenRLHF OpenRLHF 10.1k 1.0k Python 309 An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)基于 Ray 的易于使用、可扩展且高性能的 Agentic RL 框架(PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL) 2026-10-05
5 PaLM-rlhf-pytorch lucidrains 7.9k 670 Python 17 Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the PaLM architecture. Basically ChatGPT but with PaLMImplementation of RLHF (Reinforcement Learning with Human Feedback) on top of the PaLM architecture.基本上是 ChatGPT,但使用 PaLM 2026-09-20
6 InternLM InternLM 7.3k 511 Python 8 Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2025-10-30
7 Chinese-LLaMA-Alpaca-2 ymcui 7.1k 556 Python 1 中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models) 2026-04-19
8 MedicalGPT shibing624 5.9k 800 Python 6 MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型,实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。 2026-09-15
9 OpenClaw-RL Gen-Verse 5.7k 613 Python 51 OpenClaw-RL: Train any agent simply by talkingError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-05-23
10 alignment-handbook huggingface 5.7k 496 Python 92 Robust recipes to align language models with human and AI preferencesError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-09-23
11 reasoning-from-scratch rasbt 5.4k 855 Jupyter Notebook 2 Implement a reasoning LLM in PyTorch from scratch, step by step从头开始,一步步在 PyTorch 中实现推理 LLM 2026-10-02
12 transformerlab-app transformerlab 5.2k 551 Python 13 The open source research environment for AI researchers to seamlessly train, evaluate, and scale models from local hardware to GPU clusters.供 AI 研究人员无缝训练、评估和扩展从本地硬件到 GPU 集群的模型的开源研究环境。 2026-10-03
13 Kiln Kiln-AI 5.2k 392 Python 23 Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-10-09
14 argilla argilla-io 5.1k 506 Python 9 Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasetsError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-10-06
15 trlx CarperAI 4.8k 486 Python 86 A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)通过人类反馈(RLHF)进行强化学习的语言模型分布式训练的存储库 2024-01-08
16 align-anything PKU-Alignment 4.7k 503 Python 28 Align Anything: Training All-modality Model with Feedback对齐一切:通过反馈训练全模态模型 2025-11-27
17 hands-on-modern-rl walkinglabs 4.6k 334 Python 11 🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems. 🚀 一个开源的实践课程,弥合了从基本 RL 概念到 LLM 对齐、RLVR 和高级 Agentic 系统的差距。 2026-10-01
18 awesome-RLHF opendilab 4.4k 263 N/A 0 A curated list of reinforcement learning with human feedback resources (continually updated)带有人类反馈资源的强化学习精选列表(持续更新) 2026-05-20
19 ChatGLM-Efficient-Tuning hiyouga 3.7k 464 Python 0 Fine-tuning ChatGLM-6B with PEFT | 基于 PEFT 的高效 ChatGLM 微调 2023-10-12
20 docta Docta-ai 3.5k 256 Python 1 A Doctor for your data您的数据医生 2026-06-16
21 ROLL alibaba 3.4k 315 Python 102 An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language ModelsError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-10-11
22 distilabel argilla-io 3.4k 260 Python 82 Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.Distilabel 是一个合成数据和人工智能反馈框架,适用于需要基于经过验证的研究论文的快速、可靠和可扩展管道的工程师。 2026-10-06
23 Cortex qibin0506 2.7k 209 Python 0 从零构建大模型:从预训练到RLHF的完整实践 2026-08-21
24 TorchLeet Exorust 2.5k 312 Jupyter Notebook 3 LeetCode for PyTorch — 65 ML/AI interview problems from real interviews at Google, Meta, Anthropic. Jupyter notebooks, an auto-grader, and an MCP AI tutor.LeetCode for PyTorch — 65 个 ML/AI 面试问题,来自 Google、Meta、Anthropic 的真实面试。 Jupyter 笔记本、自动评分器和 MCP AI 导师。 2026-09-15
25 rlhf-book natolambert 2.4k 282 Python 1 Textbook on reinforcement learning from human feedbackError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-09-24
26 transformers_tasks HarderThenHarder 2.4k 399 Jupyter Notebook 59 ⭐️ NLP Algorithms with transformers lib. Supporting Text-Classification, Text-Generation, Information-Extraction, Text-Matching, RLHF, SFT etc.⭐️ 带有 Transformer lib 的 NLP 算法。 Supporting Text-Classification, Text-Generation, Information-Extraction, Text-Matching, RLHF, SFT etc. 2023-09-29
27 alpaca_eval tatsu-lab 2.0k 314 Jupyter Notebook 25 An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.An automatic evaluator for instruction-following language models.经过人工验证、高质量、便宜且快速。 2025-08-09
28 AgentsMeetRL thinkwee 1.9k 75 HTML 0 Awesome List for Agentic RLAgentic RL 的精彩列表 2026-10-10
29 hh-rlhf anthropics 1.9k 162 N/A 0 Human preference data for "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback"Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2025-06-17
30 ChatLM-mini-Chinese charent 1.7k 191 Python 10 中文对话0.2B小模型(ChatLM-Chinese-0.2B),开源所有数据集来源、数据清洗、tokenizer训练、模型预训练、SFT指令微调、RLHF优化等流程的全部代码。支持下游任务sft微调,给出三元组信息抽取微调示例。 2024-04-20
31 ImageReward zai-org 1.7k 90 Python 58 [NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation[NeurIPS 2023] ImageReward:学习和评估人类对文本到图像生成的偏好 2025-10-29
32 safe-rlhf PKU-Alignment 1.6k 134 Python 17 Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human FeedbackError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2025-11-24
33 WebGLM THUDM 1.6k 130 Python 51 WebGLM: An Efficient Web-enhanced Question Answering System (KDD 2023)WebGLM:高效的网络增强问答系统 (KDD 2023) 2025-03-25
34 RLHF-Reward-Modeling RLHFlow 1.5k 111 Python 21 Recipes to train reward model for RLHF.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2025-04-24
35 MOSS-RLHF OpenLMLab 1.4k 103 Python 39 Secrets of RLHF in Large Language Models Part I: PPOError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2024-03-03
36 xtreme1 xtreme1-io 1.4k 222 TypeScript 20 Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.Xtreme1 是一款用于多模态数据训练的一体化数据标记和注释平台,支持 3D LiDAR 点云、图像和 LLM。 2026-09-29
37 pyre-code whwangovo 1.3k 116 Python 4 A self-hosted ML coding practice platform. 68 problems from ReLU to flow matching — attention, training, RLHF, diffusion, and more. Instant feedback in the browser.一个自托管的机器学习编码练习平台。 68 problems from ReLU to flow matching — attention, training, RLHF, diffusion, and more.浏览器中的即时反馈。 2026-05-12
38 verl-omni verl-project 1.2k 238 Python 98 Multimodal RL training framework for diffusion & omni models用于扩散和全向模型的多模态强化学习训练框架 2026-10-10
39 labs-molt NVIDIA-NeMo 1.2k 113 Python 9 A scalable, agentic-first, and HuggingFace-native RL framework for research (9k lines). 2026-10-08
40 AI-Compass tingaicompass 976 128 Python 3 “AI-Compass”将为社区指引在 AI 技术海洋中航行的方向,无论你是初学者还是进阶开发者,都能在这里找到通往 AI 各大方向的路径。旨在帮助开发者系统性地了解 AI 的核心概念、主流技术、前沿趋势,并通过实践掌握从理论到落地的全过程。 2026-09-24
41 SimPO princeton-nlp 961 78 Python 25 [NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward[NeurIPS 2024] SimPO:具有无参考奖励的简单偏好优化 2025-02-16
42 HALOs ContextualAI 909 51 Python 7 A library with extensible implementations of DPO, KTO, PPO, ORPO, and other human-aware loss functions (HALOs).具有 DPO、KTO、PPO、ORPO 和其他人类感知损失函数 (HALO) 的可扩展实现的库。 2025-09-30
43 awesome-on-policy-distillation chrisliu298 877 38 N/A 0 A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models用于大型语言模型的策略蒸馏 (OPD) 的论文、技术报告、框架和工具的精选集 2026-10-06
44 OpenJudge agentscope-ai 871 73 Python 14 OpenJudge: A Unified Framework for Holistic Evaluation and Quality RewardsError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-09-11
45 alpaca_farm tatsu-lab 844 66 Python 6 A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data. Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2024-07-01
46 halo whitecircle 820 39 Python 12 Halo is an open-source framework built by White Circle for training large language and multimodal models 2026-10-11
47 reward-bench allenai 743 103 Python 3 RewardBench: the first evaluation tool for reward models.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-09-30
48 AlignLLMHumanSurvey GaryYufei 738 31 N/A 1 Aligning Large Language Models with Human: A Survey使大型语言模型与人类保持一致:一项调查 2026-09-08
49 Trinity-RFT agentscope-ai 708 84 Python 41 Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (LLM).Trinity-RFT 是一个通用、灵活且可扩展的框架,专为大语言模型 (LLM) 的强化微调 (RFT) 而设计。 2026-10-08
50 LLM-Algorithm-Intern-Guide Junvate 703 13 N/A 0 🚀 2026届大模型算法岗实习面经 | 包含 DeepSeek/Qwen 技术报告解析、手撕 PPO/RoPE/Transformer、RLHF 核心与八股文 | 持续更新中... 2026-03-28
51 oat sail-sg 673 63 Python 6 🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-01-29
52 Cornucopia-LLaMA-Fin-Chinese jerry1993-tech 655 66 Python 17 聚宝盆(Cornucopia): 中文金融系列开源可商用大模型,并提供一套高效轻量化的垂直领域LLM训练框架(Pretraining、SFT、RLHF、Quantize等) 2023-06-30
53 Open-AgentRL Gen-Verse 651 60 Python 10 RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic ScenariosRLAnything (ICML 2026) 和 AutoTool (ICML 2026)、DemyAgent:适用于法学硕士和代理场景的开源强化学习 2026-06-12
54 Relax redai-studio 641 177 Python 57 An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale用于大规模全模态后训练的异步强化学习引擎 2026-10-10
55 LLamaTuner jianzhnie 622 62 Python 18 Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署. 2025-01-24
56 SPPO uclaml 588 48 Python 14 The official implementation of Self-Play Preference Optimization (SPPO)Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2025-01-23
57 ARIS-in-AI-Offer wanshuiyin 586 22 Python 3 Bilingual (中文+EN) ML / LLM / diffusion / agent interview cheat sheets for AI 秋招 — generated by ARIS /interview-cheatsheet, rendered by /render-html into single-file HTML, reads anywhere — plus a CV→DBLP-fact-checked academic homepage generator and hand-authored long-form blogs 🌱双语 (中文+EN) ML / LLM / 扩散 / 代理 AI 秋招面试备忘单 — 由 ARIS /interview-cheatsheet 生成,由 /render-html 渲染为单文件 HTML,可在任何地方阅读 — 加上经过 CV→DBLP 事实检查的学术主页生成器和手工撰写的长篇博客 🌱 2026-10-06
58 Awesome-LLM-On-Policy-Distillation nick7nlp 575 16 Python 2 A curated collection of papers and resources on On-Policy Distillation for Large Language Models.有关大型语言模型的策略蒸馏的精选论文和资源集。 2026-10-06
59 TextRL voidful 565 61 Python 3 Implementation of ChatGPT RLHF (Reinforcement Learning with Human Feedback) on any generation model in huggingface's transformer (blommz-176B/bloom/gpt/bart/T5/MetaICL)Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-04-23
60 Online-RLHF RLHFlow 545 48 Python 12 A recipe for online RLHF and online iterative DPO.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2024-12-28
61 dLLM-RL Gen-Verse 523 45 Python 24 [ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA TraDo series.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-01-28
62 step_into_llm candle-org 482 127 Jupyter Notebook 27 MindSpore online courses: Step into LLMError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2025-12-22
63 LaMDA-rlhf-pytorch conceptofmind 467 73 Python 6 Open-source pre-training implementation of Google's LaMDA in PyTorch. Adding RLHF similar to ChatGPT.Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2024-02-24
64 FineEnvs adithya-s-k 463 61 Python 3 FineEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs 2026-10-08
65 LLM-RLHF-Tuning Joyce94 452 24 Python 3 LLM Tuning with PEFT (SFT+RM+PPO+DPO with LoRA) Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2023-10-11
66 ABC-GRPO chi2liu 442 42 Python 0 Code For All-Quadrant Bounded Clipping GRPO. arxiv.org/pdf/2601.03895全象限限幅裁剪 GRPO 的代码。 arxiv.org/pdf/2601.03895 2026-09-08
67 mlx-lm-lora Goekdeniz-Guelmez 429 60 Python 0 Train Large Language Models on MLX.在 MLX 上训练大型语言模型。 2026-10-05
68 VisionReward zai-org 427 15 Python 19 [AAAI 2026] VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation[AAAI 2026] VisionReward:用于图像和视频生成的细粒度多维人类偏好学习 2025-03-26
69 pykoi CambioML 410 44 Jupyter Notebook 2 pykoi: Active learning in one unified interfacepykoi:在一个统一的界面中进行主动学习 2025-09-24
70 JarvisEvo LYL1015 408 9 Python 1 [CVPR' 2026] JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator OptimizationError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2026-02-22
71 LLaVA-RLHF llava-rlhf 399 30 Python 4 Aligning LMMs with Factually Augmented RLHF将 LMM 与事实增强的 RLHF 结合起来 2023-11-01
72 quick-start-guide-to-llms sinanuozdemir 397 213 Jupyter Notebook 1 The Official Repo for "Quick Start Guide to Large Language Models"“大型语言模型快速入门指南”的官方存储库 2025-10-07
73 awesome-llm-human-preference-datasets glgh 394 19 N/A 0 A curated list of Human Preference Datasets for LLM fine-tuning, RLHF, and eval.用于 LLM 微调、RLHF 和评估的人类偏好数据集的精选列表。 2023-10-04
74 mLoRA TUDB-Labs 385 70 Python 12 An Efficient "Factory" to Build Multiple LoRA AdaptersError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2025-02-13
75 AdaRubrics alphadl 365 37 Python 1 AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent TrajectoriesAdaRubric:智能体轨迹的自适应动态评估器 2026-08-24
76 Stable-Alignment agi-templar 356 18 Python 4 Multi-agent Social Simulation + Efficient, Effective, and Stable alternative of RLHF. Code for the paper "Training Socially Aligned Language Models in Simulated Human Society".Error 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2023-06-18
77 MedQA-ChatGLM WangRongsheng 340 50 Python 3 🛰️ 基于真实医疗对话数据在ChatGLM上进行LoRA、P-Tuning V2、Freeze、RLHF等微调,我们的眼光不止于医疗问答 2023-09-02
78 ReaLHF openpsi-project 336 22 Python 0 Super-Efficient RLHF Training of LLMs with Parameter Reallocation通过参数重新分配对 LLM 进行超高效 RLHF 训练 2025-04-24
79 VADER mihirp1998 318 15 Python 11 Video Diffusion Alignment via Reward Gradients. We improve a variety of video diffusion models such as VideoCrafter, OpenSora, ModelScope and StableVideoDiffusion by finetuning them using various reward models such as HPS, PickScore, VideoMAE, VJEPA, YOLO, Aesthetics etc. 通过奖励梯度进行视频扩散对齐。我们通过使用 HPS、PickScore、VideoMAE、VJEPA、YOLO、Aesthetics 等各种奖励模型进行微调,改进了各种视频扩散模型,例如 VideoCrafter、OpenSora、ModelScope 和 StableVideoDiffusion。 2025-03-12
80 RLHF-V RLHF-V 310 8 Python 2 [CVPR'24] RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback[CVPR'24] RLHF-V:通过细粒度矫正人类反馈的行为调整,迈向值得信赖的 MLLM 2024-09-11
81 RLLoggingBoard HarderThenHarder 296 9 Python 0 A visuailzation tool to make deep understaning and easier debugging for RLHF training.一种可视化工具,可深入理解 RLHF 训练并更轻松地进行调试。 2025-02-20
82 RLHF sunzeyeah 284 35 Python 3 Implementation of Chinese ChatGPTError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2023-11-20
83 FineGrainedRLHF allenai 283 24 Python 2 2025-01-06
84 unsloth-buddy TYH-labs 281 15 Python 0 Zero-friction LLM fine-tuning skill for Claude Code, Gemini CLI & any ACP agent. Unsloth on NVIDIA · TRL+MPS/MLX on Apple Silicon. Automates env setup, LoRA training (SFT, DPO, GRPO, vision), post-hoc GRPO log diagnostics, evaluation, and export end-to-end. Part of the Gaslamp AI platform.Zero-friction LLM fine-tuning skill for Claude Code, Gemini CLI & any ACP agent. Unsloth on NVIDIA · TRL+MPS/MLX on Apple Silicon. Automates env setup, LoRA training (SFT, DPO, GRPO, vision), post-hoc GRPO log diagnostics, evaluation, and export end-to-end. Gaslamp AI 平台的一部分。 2026-06-28
85 MLX-LoRA-Studio Goekdeniz-Guelmez 277 28 Swift 0 A native Mac App for LLM fine-tuning on Apple Silicon — fully on-device, fully open source.用于在 Apple Silicon 上进行 LLM 微调的本机 Mac 应用程序 — 完全在设备上、完全开源。 2026-08-26
86 Open-R1 jianzhnie 273 55 Python 0 The open source implementation of DeepSeek-R1. 开源复现 DeepSeek-R1 2025-03-10
87 RLHF_in_notebooks ash80 254 33 Jupyter Notebook 0 RLHF (Supervised fine-tuning, reward model, and PPO) step-by-step in 3 Jupyter notebooksError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2025-06-20
88 RLHF-Label-Tool SupritYoung 254 21 Python 2 用于大模型 RLHF 进行人工数据标注排序的工具。A tool for manual response data annotation sorting in RLHF stage.用于大模型 RLHF 进行人工数据标注排序的工具。 A tool for manual response data annotation sorting in RLHF stage. 2023-08-01
89 learn-MedicalGPT bcefghj 244 14 TypeScript 1 🏥 从零基础到面试通关:20节课彻底搞懂MedicalGPT医疗大模型训练全流程 | PT/SFT/LoRA/RLHF/DPO/GRPO | 100+面试高频考点 2026-04-01
90 llama-trl jasonvanf 238 24 Python 7 LLaMA-TRL: Fine-tuning LLaMA with PPO and LoRAError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2025-08-17
91 LLaVA-MoD shufangxun 228 17 Python 3 [ICLR 2025] LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge DistillationError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2025-03-31
92 RLHF HumanSignal 227 43 Jupyter Notebook 3 Collection of links, tutorials and best practices of how to collect the data and build end-to-end RLHF system to finetune Generative AI models关于如何收集数据和构建端到端 RLHF 系统以微调生成式 AI 模型的链接、教程和最佳实践的集合 2023-07-24
93 chain-of-hindsight haoliuhl 227 16 Python 3 Simple next-token-prediction for RLHFError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2023-09-30
94 minChatGPT ethanyanjiali 226 37 Python 3 A minimum example of aligning language models with RLHF similar to ChatGPTError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2023-09-26
95 minimind-deep-dive Enping-Hu 224 14 Python 0 从 MiniMind 源码读起,再延伸到现代大模型技术体系的中文学习笔记。主线逐行精读预训练 / SFT / DPO / PPO / GRPO 与训练机制;附录 17 篇进阶卷覆盖量化、投机解码、RLHF 全景、模型代际史等 MiniMind 没涉及、但进阶绕不开的主题。 2026-07-21
96 rl-handbook lubludrova 221 13 MDX 1 A comprehensive guide to Reinforcement Learning 2026-10-02
97 Vicuna-LoRA-RLHF-PyTorch jackaduma 220 18 Python 15 A full pipeline to finetune Vicuna LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the Vicuna architecture. Basically ChatGPT but with Vicuna在消费类硬件上使用 LoRA 和 RLHF 微调 Vicuna LLM 的完整流程。在 Vicuna 架构之上实现 RLHF(带有人类反馈的强化学习)。 Basically ChatGPT but with Vicuna 2024-05-20
98 awesome-RLAIF mengdi-li 210 6 N/A 0 A continually updated list of literature on Reinforcement Learning from AI Feedback (RLAIF) 不断更新的关于人工智能反馈强化学习 (RLAIF) 的文献列表 2026-09-17
99 nanoRLHF hyunwoongko 204 19 Python 0 nanoRLHF: from-scratch journey into how LLMs and RLHF really work.nanoRLHF:从头开始了解法学硕士和 RLHF 的真正运作方式。 2026-10-07
100 VL-RLHF TideDra 204 8 Python 11 A RLHF Infrastructure for Vision-Language ModelsError 500 (Server Error)!!1500.That’s an error.There was an error. Please try again later.That’s all we know. 2024-11-15
No repositories match your search 没有匹配的仓库