Papers
arxiv:2608.01851

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

Published on Aug 3
Β· Submitted by
Aman Chadha
on Aug 7
Authors:
,
,
,
,

Abstract

Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.

Community

Paper author Paper submitter

This survey reframes modern robot learning around what actually shipsβ€”frozen policy weights versus executable skillsβ€”and introduces a five-rung taxonomy of code-as-policy systems based on increasingly powerful combinations of execution feedback, persistent memory, and program search, culminating in autonomous self-improving robot skill loops.

➑️ 𝐊𝐞𝐲 𝐇𝐒𝐠𝐑π₯𝐒𝐠𝐑𝐭𝐬 𝐨𝐟 𝐭𝐑𝐞 π–πžπ’π π‘π­π¬-𝐯𝐬-𝐒𝐀𝐒π₯π₯𝐬 π…π«πšπ¦πžπ°π¨π«π€:

🧭 π‘Ύπ’†π’Šπ’ˆπ’‰π’•π’” 𝒗𝒔. π‘Ίπ’Œπ’Šπ’π’π’” π‘»π’‚π’™π’π’π’π’Žπ’š: Introduces a unified taxonomy spanning 77 core systems across six robot-learning familiesβ€”code-as-policy, end-to-end VLA, reward synthesis, skill libraries, sim-to-real/transfer, and benchmarksβ€”plus 225 landscape works. Its key architectural distinction is whether competence is encoded in frozen neural weights (e.g., VLA backbone + action head) or represented as inspectable, executable programs/skills that can be edited and recombined after deployment. The taxonomy on page 3 makes this decomposition explicit.

πŸ”„ π‘­π’Šπ’—π’†-π‘Ήπ’–π’π’ˆ 𝑺𝒆𝒍𝒇-π‘°π’Žπ’‘π’“π’π’—π’†π’Žπ’†π’π’• 𝑳𝒂𝒅𝒅𝒆𝒓 (𝑭 + 𝑴 + 𝑺): The paper's main analytical novelty is decomposing code-as-policy agents by three operational mechanismsβ€”Feedback (F) from execution, Memory (M) persisted across tasks, and Search (S) over multiple candidate programsβ€”and arranging systems from zero-shot synthesis β†’ closed-loop repair β†’ skill-library accumulation β†’ evolutionary search β†’ full F+M+S self-improvement. Crucially, it distinguishes sequential debugging from genuine search and frozen model parameters from runtime memory, making β€œself-improvement” technically testable rather than a loose label.

🧠 𝑭𝒖𝒍𝒍 𝑺𝒆𝒍𝒇-π‘°π’Žπ’‘π’“π’π’—π’Šπ’π’ˆ 𝑹𝒐𝒃𝒐𝒕 𝑳𝒐𝒐𝒑 + π‘Ίπ’Œπ’Šπ’π’ π‘¬π’„π’π’π’π’Žπ’š: Identifies a sparsely populated frontierβ€”represented by ASPIRE, ENPIRE, and RoboClawβ€”where an agent executes skills, obtains grounded traces/feedback, stores validated skills in persistent memory, and searches/mutates candidate programs, feeding accumulated competence into future tasks. The architecture is summarized in the page-11 diagram as Actor Agent β†’ Execution Engine (F) β†’ Skill Memory (M) β†’ Evolutionary Search (S) with skills recursively returned to future tasks. The survey argues this is the missing adaptation layer between today's static robot-skill marketplaces and genuinely deployable, continually improving robot ecosystems.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.01851 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.01851 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.01851 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.