slime-rl-training Skill
Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling. Published by NousResearch in hermes-agent.
What is slime-rl-training Skill?
Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling. Published by NousResearch in hermes-agent. This profile combines repository metadata with install, compatibility, and usage signals so developers can quickly decide whether it fits their agent workflow before opening the source repository.
Automated repository signals based on public metadata such as recency, license, installation evidence, and adoption. These are not a security audit or endorsement.
Key capabilities
- Includes SKILL.md support
- Reusable instructions support
- Data analysis
- Research
- Data analysis use cases
- Research use cases
Technical details
- Install or run with Copy skill directory
When to use slime-rl-training Skill
- Use it for data analysis.
- Use it for research.
Built with
Editorial notes
Source
- Creator: NousResearch
- Repository: NousResearch/hermes-agent
- Skill file: optional-skills/mlops/slime/SKILL.md
What it does
Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling.
Skill instructions
slime: LLM Post-Training Framework for RL Scaling slime is an LLM post-training framework from Tsinghua's THUDM team, powering GLM-4.5, GLM-4.6, and GLM-4.7. It connects Megatron-LM for training with SGLang for high-throughput rollout generation. When to Use slime Choose slime when you need: - Megatron-LM native training with SGLang inference - Custom data generation workflows with flexible data buffers - Training GLM, Qwen3, DeepSeek V3, or Llama 3 models - Research-grade framework with production backing (Z.ai) Consider alternatives when: - You need enterprise-grade stability features → use miles - You want flexible backend swapping → use verl - You need PyTorch-native abstractions → use torchforge Key Features - Training: Megatron-LM with full parallelism support (TP, PP, DP, SP) - Rollout: SGLang-based high-throughput generation with router - Data Buffer: Flexible prompt management and sample storage - Models: GLM-4.x, Qwen3, DeepSeek V3/R1, Llama 3 Architecture Overview ┌─────────
Explore related resources
Frequently asked questions
What is slime-rl-training?
slime-rl-training is a open-source AI agent skill with Copy skill directory. Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework.
Who is slime-rl-training best for?
slime-rl-training is best for reusing agent instructions, scripts, and references, data analysis workflows, research workflows.
How do I install slime-rl-training?
Install or run slime-rl-training using Copy skill directory. Check slime-rl-training for the latest setup command.
Is slime-rl-training actively maintained?
slime-rl-training may need a closer maintenance check before production use.
Auto-fetched from GitHub.