SYNTHESIS NOTE
Topics›Domain Specialization›this note

Can prompt optimization teach models knowledge they lack?

Explores whether sophisticated prompting techniques can inject new domain knowledge into language models, or if they're limited to activating existing training knowledge.

Synthesis note · 2026-02-21 · sourced from Domain Specialization

The knowledge injection survey makes this constraint explicit: prompt optimization "focuses on fully leveraging or guiding the LLM to utilize its internal, pre-existing knowledge." It does not retrieve from external sources. It does not update parameters. It works entirely within the model's existing knowledge distribution.

This is a hard ceiling, not a soft limitation. When a domain requires knowledge that the model was never trained on — proprietary documents, post-training regulations, specialized ontologies, organization-specific processes — no prompting strategy can supply it. The model can reorganize, foreground, or combine what it knows, but it cannot know what it was never trained to know.

The practical consequence shows up in two failure modes. First, models prompted to act as domain experts will confidently apply general-purpose reasoning patterns to domain-specific problems where those patterns don't hold. The prompt activates "medical reasoning" as a behavioral style, not as medical knowledge. Second, prompt performance depends on how thoroughly the domain is represented in pre-training — well-documented domains (clinical guidelines, legal statutes, financial regulations) are more promptable than proprietary or emerging domains.

This makes prompt-only domain specialization a form of retrieval from fixed memory. The memory can be searched more or less skillfully, but it can't be expanded. Every sophisticated prompting technique — few-shot examples, chain-of-thought elicitation, role specification — is fundamentally retrieval from training data, dressed as reasoning.

The implication is that the right question before choosing prompt optimization is not "how should we phrase this prompt?" but "is the required domain knowledge in the model's training distribution?" If yes, prompting is sufficient and efficient. If no, the investment must go into a different injection paradigm — dynamic retrieval, fine-tuning, or adapter layers.

Since Why do specialized models fail outside their domain?, there's a version of the ceiling problem in the opposite direction: models that are fully fine-tuned can know a domain deeply while losing general coverage. Prompt optimization avoids the cliff problem by not modifying parameters — but only by accepting the ceiling problem instead. Every approach involves a trade-off; this one chooses breadth over depth.

Reynolds & McDonell (2021) provide the upstream mechanism: few-shot prompting is "task location in the model's existing space of learned tasks" — not task learning. Alternative 0-shot prompts that communicate task intention through natural language semiotics match or exceed few-shot performance, confirming that the model already has the capability and the prompt's job is to locate it. Meta-prompt programming further extends this: the LLM itself can be prompted to write task-specific prompts, offloading the location search to the model's own understanding of its capabilities.

Inquiring lines that read this note 219

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does AI assistance promote real skill development or substitute for independent learning? How do prompting refinements mask underlying biases and model frequency patterns? How should conversational recommenders balance preference elicitation with direct recommendation? Can prompt-based context override biases that were embedded during pretraining? What reasoning architectures enable models to solve complex problems efficiently? Why does adding new knowledge through fine-tuning degrade existing capabilities? Why do language models resist personality conditioning through prompts? Why do LLM recommenders underperform collaborative filtering despite their capabilities? What training dynamics and scale trigger emergence of reasoning capabilities? Why do token-level mechanisms matter for learning to reason? Why can't prompting alone inject genuinely new knowledge into models? What capability trade-offs arise from domain specialization through fine-tuning? Can models improve accuracy without degrading reasoning quality? How do prompt design choices influence model reasoning and performance? Does encoded knowledge in language models actually influence their outputs? How should retrieval systems handle complex multi-step reasoning? Is reasoning capability latent in base models or created by post-training? What makes distillation transfer some model capabilities while suppressing others? What prevents conversational agents from taking initiative in dialogue? What compositional reasoning failures limit large language models despite scale? When do semantic similarity approaches miss structural retrieval failures? Why don't LLMs reliably translate capability into accurate outputs? Does RL create genuinely new reasoning capabilities or refine existing ones? Why do stronger reasoning capabilities create tradeoffs with instruction following? How much do training data properties shape model reasoning? Can inoculation prompting prevent emergent misalignment after reward hacking? What causes retrieval-augmented generation systems to fail despite access to external knowledge? Is language model reasoning authentic and what causes models to reason? Do language models learn genuine understanding or just surface patterns? How does improved reasoning affect models' ability to acknowledge uncertainty? Can diffusion models match autoregressive performance on language generation tasks? How should inference compute be allocated based on problem difficulty? Do reasoning benchmarks predict model performance in long-horizon workflows? How do standardized protocols improve multi-agent coordination and reliability? Does model confidence reliably signal actual accuracy in practice? How effectively can language models perform reasoning, especially combined with symbolic methods? How should systems decide whether to retrieve or reason alone? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? What articulatory and acoustic information does speech preserve that transcription destroys? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? What factors drive AI persuasiveness and how can it be mitigated? What causes reasoning models to fail or wander off track? Can mechanistic interpretability reliably guide practical model design choices? Should agents decouple planning from perception grounding for better performance? What training data selection strategies maximize generalization across difficulty levels? What types of diversity prevent reasoning systems from collapsing? How does synthetic data quality and diversity affect downstream model capabilities? Can language models build genuine grounding through interaction? How much does training format versus domain influence reasoning? When should work require human-AI partnership versus full automation? Do language models reason like humans or mimic surface patterns?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
23 direct connections · 248 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

prompt optimization cannot inject new knowledge — it can only activate knowledge the model already contains