Can language models learn skills without human supervision?
Can a three-role self-play system—Challenger, Reasoner, Judge—bootstrap natural-language skills from raw context alone, without human labels or external reward signals?
Ctx2Skill closes the skill-construction loop without human annotation or an external reward signal by running a three-role self-play loop. A Challenger generates probing tasks and rubrics against a context; a Reasoner attempts them guided by its current skill set; a neutral Judge issues binary pass/fail feedback. The signal is internal — easily-solved tasks are routed back to strengthen the Challenger, while failed cases are routed to Proposer and Generator agents that synthesize targeted skill updates for the Reasoner. Both sides evolve through accumulated natural-language skills rather than parameter updates.
This matters because it dissolves the two bottlenecks that block automated skill construction: the prohibitive cost of manually annotating skills for long, dense contexts, and the absence of external feedback to tell automated construction what to improve. Self-play manufactures the missing feedback — the Challenger's escalating difficulty is the curriculum, and the Judge's binary verdict is the reward — so the system bootstraps a skill set for an arbitrary context from nothing but the context itself.
The counterpoint, which the paper takes seriously, is adversarial collapse: a Challenger free to maximize difficulty drifts toward extreme tasks, and a Reasoner chasing them accumulates over-specialized skills that no longer generalize. Self-play that only ratchets pressure destroys itself. This is why Ctx2Skill needs a separate replay mechanism to anchor generality — which is the tension worth tracking. Therefore the insight is real but conditional: unsupervised co-evolution of language skills works only when adversarial pressure is balanced against a generalization safeguard.
Inquiring lines that read this note 57
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can we prevent synthetic data from contaminating statistical inference and corpora? Does AI assistance promote real skill development or substitute for independent learning?- Why do Generation-Then-Comprehension and AI Delegation produce opposite learning outcomes?
- Does AI assistance transfer learning gains to independent tasks without scaffolding?
- Why does online RL succeed where supervised training fails for self-correction?
- Why does asymmetric self-play create naturally calibrated difficulty better than fixed curricula?
- Can synthetic self-play data teach models when to disagree?
- How do instruction backtranslation and MAGPIE demonstrate self-generation principles?
- How does adversarial collapse threaten unsupervised self-play skill construction?
- What makes self-consistency a sufficient training target for the judge role?
- What role does natural language play in breaking reinforcement learning performance plateaus?
- Can verifier-free RL work without manual preference labels or task-specific training?
- How can verifier-free reinforcement learning handle reasoning without task-specific checks?
- Why does natural language feedback break performance plateaus that numerical rewards alone cannot?
- How do graduated phase rewards emerge complex dialogue behavior from simple objectives?
- Can structured natural language feedback outperform scalar rewards in RL?
- Can AI learn intrinsic motivation to assess its own relevance?
- How does hidden processing in language models prevent accurate self-assessment?
- Can we separate task competence from genuine agency in language model outputs?
- Can humans learn accurate models of AI through repeated interaction without labels?
- Can metacognitive categories be learned instead of fixed by human designers?
- Can subjective tasks be delegated without human feedback loops?
- Does human-in-the-loop AI collaboration accelerate recursive self-improvement safely?
- What causes gradient-based steering via natural language descriptions to work?
- How tight should a textual learning rate be before it prevents skill escape?
- Where does skill extraction fail compared to genuine model adaptation?
- How can language models extract more value from fewer demonstrations?
- Can self-supervised methods replace human annotations for process reward models?
- Can self-supervised process models replace human annotations at scale?
- Does self-supervised process supervision work for domains with ambiguous correctness?
- Can trajectory structure alone provide process supervision without human annotation?
- What role does self-learning play in improving agent reasoning without annotation?
- Does self-play feedback improve skills created from the agent's own experience?
- Can AI systems improve themselves without external feedback?
- How can agents evolve their own skills without human input?
- Can self-improving agents become truly autonomous without intrinsic metacognition?
- Can binary judge feedback replace external reward signals for skill learning?
- How do self-play and human-anchored rewards separate competence from convention?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can language models improve themselves without any external training data?
Explores whether two language models playing against each other—one generating questions, one solving them—can create a self-improving loop. Matters because it would eliminate dependence on human-labeled datasets.
same proposer-vs-solver self-play structure; Ctx2Skill adds a third neutral Judge role and evolves natural-language skills rather than model weights
-
Does creating skills inside the agent loop eliminate mismatches?
Can coupling skill creation directly to the runtime reasoning loop—rather than authoring skills offline—close the gap between when skills are made and when they're used? This matters for whether agents can ground new capabilities in their actual situated context.
both ground skill creation in the agent's own experience; Ctx2Skill manufactures the missing feedback via self-play where MUSE manufactures it via in-loop invocation
-
Can skill documents be optimized like neural network weights?
Explores whether natural-language skill artifacts—packaging procedures, heuristics, and policies—can be systematically improved through iterative editing and validation, similar to how gradient descent refines model parameters.
shares the generalization-collapse risk: Ctx2Skill needs a replay safeguard against adversarial drift much as SkillOpt needs a held-out gate to prevent overfitting self-edits
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- From Context to Skills: Can Language Models Learn from Context Skillfully?
- SPICE: Self-Play In Corpus Environments Improves Reasoning
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Self-Questioning Language Models
- Training Language Models to Self-Correct via Reinforcement Learning
- Self-Rewarding Language Models
- PretrainZero: Reinforcement Active Pretraining
Original note title
challenger-reasoner-judge self-play can co-evolve natural-language skills with no human supervision