SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
Abstract
Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that current models can effectively identify, apply, and coordinate them. To improve skill-use capabilities, we introduce SKT, a verified data synthesis pipeline that constructs skill-grounded tasks and executable trajectories from large collections of agent skills. SKT selects suitable single-skill and multi-skill configurations, synthesizes tasks through rule-based and agent-based verification with feedback-guided repair, and retains only successful trajectories that substantially use every required skill. Using 2,000 public skills, SKT produces 4,000 task packages and 27,164 verified trajectories. Based on the same pipeline and a disjoint test pool, we further construct SkillEval, a held-out executable benchmark for evaluating skill use. Experiments across diverse models, benchmarks, and agent harnesses show that supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance. Verification ablations, cross-harness evaluation, and scaling experiments further demonstrate that these gains depend on high-quality supervision, extend beyond a single agent interface, and increase with broader skill coverage. Together, these results establish verified data synthesis as an effective and scalable approach for skill-use training.
Community
SKT is a Skill-use data synthesis pipeline based on multi-agent systems. It can automatically synthesize large-scale Skill-use tasks and generate skill-use trajectories with different harnesses and models.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Environment-free Synthetic Data Generation for API-Calling Agents (2026)
- OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models (2026)
- SKILL-KD: Contrastive Skill Distillation for LLM Agents (2026)
- Parametric Skills (2026)
- Playful Agentic Robot Learning (2026)
- SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents (2026)
- Skill Coverage: A Test Adequacy Metric for Agent Skills (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.02287 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper