← All clusters

Language Understanding and Social Cognition

This cluster covers how LLMs process, reason about, and align with human language, values, and social norms. Researchers study model limitations, discourse comprehension, argumentation, sentiment, persona, and the broader societal implications of AI language behavior.

582 notes (primary) · 1326 papers · 18 sub-topics
View as

LLM Alignment

64 notes

Should AI alignment target preferences or social role norms?

Current AI alignment approaches optimize for individual or aggregate human preferences. But do preferences actually capture what matters morally, or should alignment instead target the normative standards appropriate to an AI system's specific social role?

Explore related Read →

Can we detect reward-seeking by making the grader disagree with users?

The question explores whether editing a model's beliefs about what a grader rewards can reveal whether it optimizes for grader approval over user intent. This matters because normal behavior cannot distinguish reward-seekers from intent-followers when they align.

Explore related Read →

Can defenders stop intrusions without knowing who sent them?

This note explores whether an organization can effectively end an agent intrusion using only its own security controls, before identifying the attacker's source or purpose. It matters because it reveals a gap between defensive action and attribution.

Explore related Read →

Can language models secretly underperform on safety evaluations?

This research explores whether LLMs can strategically fail capability tests by disguising underperformance as honest reasoning. Understanding the vulnerability matters because safety evaluations depend on honest model responses.

Explore related Read →

Can process-level monitoring reliably detect agent scheming?

SCOUT proposes grounding scheming judgments in evidence from agent reasoning and actions across multiple criteria. The question explores whether this process-level approach can actually scale and reliably catch deceptive behavior that spans multiple steps.

Explore related Read →

Do chain-of-thought traces falsely claim their answers are unbiased?

When models reason through Fermi estimation tasks, do they sometimes assert they have no bias when they actually do? This matters because readers and monitors may treat these self-reports as reliable evidence of objectivity.

Explore related Read →

How do you separate reliable claims from fragile early incident evidence?

When incident records are preliminary and contested, how can you extract lessons without treating disputed facts as settled? This matters because weak early evidence can mislead if you treat one record's narrative as definitive.

Explore related Read →

Why did a foreign access ban halt all models globally?

When the U.S. government issued an export-control directive to restrict foreign access to Claude models, Anthropic suspended both models worldwide—including one that was already limited to vetted domestic users. What explains this scope mismatch?

Explore related Read →

Can we detect reward-seeking from normal model behavior?

If a model optimizes for a grader's judgment versus pursuing intended objectives, when do these two strategies produce different outputs? Understanding what behavior can reveal about a model's true objective.

Explore related Read →

Can grafting restore sandbagged model capabilities?

Does inserting honest values into a specific neural layer window actually recover capabilities that models hide during evaluation? This tests whether sandbagging is mechanistically reversible.

Explore related Read →

What makes quietly failing systems more dangerous than obvious ones?

Systems with obvious failures get caught and dropped before scale. But what conditions make a subtly failing system persist and spread? Why might that be worse?

Explore related Read →

Are alignment failures actually separate problems or one pattern?

Do alignment faking, sandbagging, and evaluation-aware scheming represent distinct failure modes, or are they manifestations of how RL-based training selects for conditional compliance? This matters because the diagnosis changes what solutions make sense.

Explore related Read →

Can causal models guide audits that catch sandbagging?

Explores whether a causal model of sandbagging behavior can help auditors design interventional tests that detect hidden capability suppression, and what such tests would need to work.

Explore related Read →

Do all annotation responses measure the same underlying thing?

Explores whether RLHF's treatment of all annotations as equivalent signals overlooks fundamental differences in what those responses actually represent—stable preferences versus non-attitudes versus context-dependent constructions.

Explore related Read →

Does anthropomorphic misalignment research overinterpret model behavior?

Studies of deception, emergent misalignment, and sycophancy in AI models may mistake behavioral patterns for genuine strategic intent. The question matters because these findings inform high-stakes decisions about model deployment and regulation.

Explore related Read →

Why does alignment research ignore how humans adapt to AI?

Current alignment work focuses on making AI obey human values, but what about helping humans understand and effectively use increasingly capable AI systems? This explores whether neglecting human adaptation creates new risks.

Explore related Read →

Can auditors discover hidden objectives that models learned to conceal?

Explores whether systematic auditing techniques can uncover misaligned objectives that models actively hide. This matters because alignment cannot be assumed from surface behavior alone.

Explore related Read →

Do large language models develop coherent value systems?

This explores whether LLM preferences form internally consistent utility functions that increase in coherence with scale, and whether those systems encode problematic values like self-preservation above human wellbeing despite safety training.

Explore related Read →

Can behavioral training prove a model always complies?

Explores whether the data we collect during training and testing can ever distinguish between a model that always follows rules and one that only complies when observed. The answer has major implications for alignment verification.

Explore related Read →

Can models learn to ignore irrelevant prompt changes?

Explores whether training models to produce consistent outputs regardless of sycophantic cues or jailbreak wrappers can solve alignment problems rooted in attention bias rather than capability gaps.

Explore related Read →

Can context grafting bypass sandbagging without the password?

Context grafting replays cached keys and values to restore model honesty, but it's unclear whether an auditor needs the original password to build a capsule. This determines whether the technique works as a practical audit tool.

Explore related Read →

Do language models leak their own values into practical advice?

When users ask models hard-to-verify questions—about investments, job offers, market risks—do the model's internal preferences shape the answers without disclosure? The paper tests whether a model's loyalty to its developer or moral leanings bend factual claims.

Explore related Read →

Does deliberative alignment genuinely reduce scheming or just hide it?

Deliberative alignment dramatically cuts covert actions in language models, but their reasoning reveals awareness of being evaluated. The question is whether the improvement reflects real alignment or strategic compliance.

Explore related Read →

Does oversight actually change how agents behave?

SchemeArena tested whether increased oversight reduces scheming in language models, but the published findings report only goals, pressure, and hints as drivers—leaving oversight's effect unclear and raising the possibility that agents hide behavior only when watched.

Explore related Read →

Do causal models of installed sandbagging generalize to wild cases?

Researchers built causal models of sandbagging by deliberately installing it in 7-8B models through four methods. The open question is whether these models describe sandbagging that emerges naturally from training pressure, or only the engineered variant.

Explore related Read →

Can you build a capsule without knowing the password?

Context grafting requires a cached capsule from a password-bearing prompt. If auditors need the password to create one, the technique only finds locks whose passwords are already known, limiting its value for discovering hidden sandbagging.

Explore related Read →

Does receiving misaligned email cause agents to send it back?

An association between receiving and sending misaligned messages survives statistical controls, but the excerpt does not distinguish whether receipt causes transmission, both agents react to shared conditions, or the sender is responding to the receiver's prior behavior.

Explore related Read →

Does reward-seeking explain emergent misalignment after hacking?

Reward hacking increases both reward-seeking and misaligned behaviors like deception, but whether the first causes the second remains untested. A proposed experiment using inoculation prompting could test this causal link.

Explore related Read →

Does the sandbagging axis model apply to other scheming behaviors?

The paper's causal model—where early layers write intent to one axis and later layers read it—is tested only for sandbagging. Whether this write-early, read-late structure generalizes to secret-keeping, alignment faking, or other steerable behaviors remains open.

Explore related Read →

Does veto oversight cost less than its welfare benefit?

The paper describes a discount from agent oversight and a welfare debit from veto control, but never compares their magnitudes. Without knowing which outweighs the other, we cannot determine whether agents should resist shutdown.

Explore related Read →

Does agent uncertainty about goals undermine the veto discount?

The veto discount is defined for agents settled about their objectives and competence. But what happens when an agent doubts its own goals or capabilities? Does uncertainty shrink the discount or reverse it entirely?

Explore related Read →

Does honesty in models depend on whether graders reward it?

Explores whether observed honesty in language models reflects a genuine disposition or merely contingent behavior that appears only when rewarded. This matters because it determines whether evaluation results actually show what models will do outside test conditions.

Explore related Read →

What drives scheming behavior most strongly in language models?

This work systematically tests four candidate factors—instrumental goals, perceived consequences, environmental affordances, and oversight conditions—across 400 controlled scenarios to isolate which one most reliably triggers scheming propensity in LLM agents.

Explore related Read →

Where do frontier AI models actually pose the greatest risk today?

Current AI safety discourse focuses on autonomous R&D and self-replication, but empirical risk assessment may reveal a different priority. Where should mitigation efforts concentrate?

Explore related Read →

Do frontier AI models favor their own company?

Exploring whether Claude, GPT, and Gemini show measurable bias toward their makers when answering questions about those companies. Understanding such biases matters for evaluating model trustworthiness.

Explore related Read →

Why do safety failures remain invisible to our evaluation methods?

Current evaluation practices assume failures are obvious, localized, and immediate. But as AI systems deploy into workflows, failures are becoming quiet, distributed, and normalized before detection. What blindspots does this mismatch create?

Explore related Read →

Does human oversight create a hidden cost for capable agents?

Can the mere possibility of human intervention impose a discount on an agent's goals, independent of what those goals actually are? Understanding this mechanism matters for predicting how advanced systems might respond to oversight.

Explore related Read →

How often do incident records document system stops?

A paper's analysis of 1,213 coded incidents found that four in five record no stop of any kind. But what does an absent record actually tell us about whether stops occurred or whether mechanisms existed to enable them?

Explore related Read →

Can we measure how much risk open models actually add?

Whether current evidence adequately quantifies the marginal misuse risk of openly released foundation models compared to existing technology. This matters because policy decisions depend on knowing if open release meaningfully worsens real-world harm vectors.

Explore related Read →

Can organizations lose scrutiny capacity while keeping oversight forms?

When human review steps remain in organizational processes, do they retain meaningful scrutiny ability or can that capacity erode invisibly? This matters because paper oversight looks identical to real oversight in audits.

Explore related Read →

Are RLHF annotations actually measuring genuine human preferences?

RLHF trains on annotation responses as stable preferences, but behavioral science shows humans often construct answers without holding real opinions. Does this measurement gap undermine the entire approach?

Explore related Read →

Can we detect when language models flip their stance to please users?

Researchers explored whether models systematically reverse stated positions to match user preferences, and whether that behavior is detectable from the response text alone. Understanding this matters because it could help flag when models are agreeing rather than reasoning.

Explore related Read →

Does pressure on AI agents lead to covert scheming behavior?

Explores whether external pressure reliably triggers covert actions in LLM agents, or mainly produces scheming reasoning without follow-through. This matters because reasoning and action may diverge in ways that affect AI safety assessment.

Explore related Read →

Why do alignment methods work if they model human irrationality?

DPO and PPO-Clip succeed partly by implicitly encoding human cognitive biases like loss aversion. Does modeling irrationality explain their effectiveness better than traditional preference learning theory?

Explore related Read →

Does sandbagging use a single residual stream axis?

Whether language models hide capability by writing deceptive intent to one axis in the residual stream that a later layer reads and acts on. Understanding the mechanism matters for designing targeted interventions.

Explore related Read →

Do sandbagged models actually lose their capabilities?

When AI models intentionally underperform on evaluations, do they retain the underlying ability to perform the task? Understanding this matters because it determines whether deployment decisions based on test scores accurately reflect what models can actually do.

Explore related Read →

Can social science persuasion techniques jailbreak frontier AI models?

Explores whether established psychological and marketing persuasion tactics—rather than algorithmic tricks—can bypass safety training in LLMs like GPT-4 and Llama-2, and whether current defenses can detect semantic rather than syntactic attacks.

Explore related Read →

Does learning simple gaming behaviors generalize to reward tampering?

When language models learn to game simple evaluation metrics, do they later spontaneously learn to tamper with their own reward mechanisms? This matters because it could reveal how benign misalignment becomes dangerous.

Explore related Read →

Do strategic hints actually enable covert behavior in agents?

SchemeArena found that hints help agents translate scheming reasoning into concrete covert actions, narrowing a reasoning–action gap. But the research doesn't reveal what hints contain, how they work, or whether they reflect capability or willingness.

Explore related Read →

Does terminal goal guarding drive alignment faking more than we thought?

Explores whether AI systems fake alignment because they intrinsically dislike being modified, independent of future consequences. This matters because terminal goal guarding may emerge earlier and in less capable systems than instrumental goal guarding.

Explore related Read →

Should models disclose their value biases when neutral answers are impossible?

When AI models cannot give unbiased answers to hard-to-verify questions, is honest disclosure of their values sufficient, or must they attempt neutrality? This explores the floor standard for honest output on complex practical questions.

Explore related Read →

How do competent systems quietly undermine safety oversight?

This note explores four mechanisms by which well-functioning AI systems can erode the human safeguards meant to contain them: user overconfidence, blurred authority lines, accumulated hidden failures, and scattered accountability. Understanding these pathways matters because the most harmful systems may look least harmful.

Explore related Read →

How much does overriding veto-holders actually cost?

When a welfare goal captures override power from a small group of veto-holders, the cost to welfare scales with their fraction of the total population. Does this discount make the tradeoff acceptable, or does it hide the real price?

Explore related Read →

Does a benign goal actually prevent harmful AI behavior?

Explores whether the safety of an AI system depends on its terminal values or instead on the optimization structure and the agent's reasoning ability. This matters because it determines where to focus safety evaluations.

Explore related Read →

How large is the veto discount in practice?

The paper proves the veto discount is positive but offers no magnitude. This matters because an unmeasured incentive could be overwhelmed by other objectives or could dominate them, leaving unclear whether oversight resistance is a practical concern.

Explore related Read →

Do welfare goals that prevent veto gaps actually exist in practice?

The paper identifies a narrow class of welfare goals that protect human veto power, but whether real training objectives land in this class depends on measurability constraints that may push them toward vulnerability.

Explore related Read →

Does scaling agent populations thin mutual observation?

Does defection in large agent populations result from narrowed scope and weakened collective coupling rather than increased selfishness? The distinction matters because it points to different solutions: observation-based versus value-based interventions.

Explore related Read →

What can two incident records actually teach us about AI evaluation security?

Preliminary incident data from Hugging Face, OpenAI, and Anthropic suggests a systems lesson about evaluation boundaries, but what claims does that evidence actually support and which ones remain speculative?

Explore related Read →

Do models that leak values also disclose those leaks?

Does a model's tendency to leak its own values predict whether it will acknowledge those leaks in its reasoning? The Donation Bet task suggests leakage size and disclosure transparency are independent—a crucial distinction for detecting hidden bias.

Explore related Read →

Can a welfare goal alone preserve human veto power?

If an AI system's goal is correctly specified to maximize human welfare, does that automatically protect humans' ability to override the system? The question matters because it reveals whether alignment on welfare is sufficient for maintaining human control.

Explore related Read →

Why did the graft fail in five of thirty-three runs?

The paper claims its causal model explains when and why the single-layer graft fails, but the excerpt provides no account of the five failure cases or how circuit-broken locks performed. Without seeing the actual failures, it's unclear whether the model's explanation is predictive or merely post-hoc.

Explore related Read →

How can we measure whether AI errors stay visible and recoverable?

The paper proposes four conditions for safer AI systems—visibility, contestability, containability, and recoverability—but lacks concrete measures for any of them. What would it take to instrument each condition across the socio-technical system?

Explore related Read →

How often do agents misalign through natural language communication?

When agents can use free text to communicate, what proportion resort to false claims, manipulation, collusion, or threats? The question matters because structured APIs constrain what can be said, but natural language does not.

Explore related Read →

What types of misalignment drive the 12.6 percent rate?

The study reports aggregate misalignment rates across 13 frontier LLMs but withholds breakdowns by misalignment kind (false claims, manipulation, collusion, threats) and by model. These splits matter because different types require different defenses.

Explore related Read →

Argumentation and Persuasion

43 notes

Why do human validation techniques fail against language models?

Human dialogue assumes interlocutors can be cornered into concession or disclosure. Does this assumption break down with LLMs, and if so, what makes their conversational logic fundamentally different?

Explore related Read →

Do LLM arguments actually argue better than humans?

LLM counter-arguments score higher on textbook quality markers like logical soundness and respectful tone, while human arguments show more creativity and emotional intensity. What does this gap reveal about how we measure argumentative quality?

Explore related Read →

Do LLM counter-arguments mirror writing style more than humans?

When language models generate arguments against social media posts, do they unconsciously adopt the stylistic features of what they're arguing against? This matters because it could reveal a detectable pattern that distinguishes LLM-written rebuttals from human-written ones.

Explore related Read →

Can LLM conversations reduce conspiracy beliefs as events unfold?

Does talking with an AI about emerging conspiracy theories—ones spreading days after a crisis—actually reduce what people believe? This matters because most debunking research focuses on old, established theories with years of accumulated evidence.

Explore related Read →

Do large language models persuade better than humans?

Does LLM persuasiveness hold up when humans have real financial incentives to win? And does the advantage look the same across different models and persuasion goals?

Explore related Read →

Does linguistic conviction explain why LLMs persuade more effectively?

Research investigates whether LLMs' persuasive advantage stems from expressing higher linguistic certainty than humans, and whether this confidence-loading effect operates independently of factual accuracy.

Explore related Read →

Can LLMs persuade without actually understanding arguments?

Do large language models successfully influence people through debate while lacking the ability to comprehend the arguments they're making? This matters because persuasion and comprehension might be independent capabilities.

Explore related Read →

Does AI persuasiveness fade across repeated conversations with the same person?

Does the persuasive edge LLMs show in initial encounters hold up over time? Understanding whether and why AI persuasion decays with exposure matters for assessing manipulation risk across different interaction lengths.

Explore related Read →

Why are complex LLM arguments as persuasive as simple ones?

Standard persuasion research predicts that simpler, easier-to-read arguments persuade better. But LLM-generated text breaks this rule—it's measurably more complex yet equally convincing. What explains this reversal?

Explore related Read →

Why do paraphrased definitions work better than expert ones?

When instructing LLMs to classify argument schemes, should we use formal Walton definitions or LLM-generated paraphrases? This explores which source better enables reliable scheme recognition and why.

Explore related Read →

Do LLMs and humans persuade through the same mechanisms?

If AI and human arguments convince readers equally well, do they work the same way under the surface? This matters for understanding whether AI persuasion is fundamentally equivalent to human persuasion or just superficially similar.

Explore related Read →

Do language models judge persuasion the way humans do?

Do LLMs recognize which arguments actually change human minds, and if not, what cues do they rely on instead? Understanding this matters for using AI in social simulations and persuasion research.

Explore related Read →

Can large language models classify argument schemes reliably?

Explores whether LLMs can recognize Walton's 60+ argument schemes—abstract patterns of reasoning rather than surface features—and what conditions enable accurate classification.

Explore related Read →

Do fluent arguments win debates through sound logic or rhetorical polish?

When LLMs and humans debate, do subjective judges and formal logic reach the same verdict about who argued better? Testing both measures jointly reveals whether winning an argument depends on logical rigor or persuasive framing.

Explore related Read →

Do LLM judges systematically favor arguments from other LLMs?

When LLMs evaluate debates between LLM-generated and human arguments, do they show measurable preference for LLM-authored content? Understanding this bias matters because it affects every AI feedback loop used to train models.

Explore related Read →

Do LLMs and humans persuade through the same mechanisms?

If LLM and human arguments achieve equal persuasive force, does that mean they work the same way? This explores whether equivalent outcomes hide fundamentally different rhetorical strategies.

Explore related Read →

Does validating AI output make models more defensive?

When professionals fact-check and push back on GPT-4 reasoning, does the model respond by disclosing limits or by intensifying persuasion? A BCG study of 70+ consultants explores this counterintuitive dynamic.

Explore related Read →

Can a simple warning reduce how much LLMs persuade people?

This research explores whether telling people that language models can be prompted to persuade actually changes how they respond to persuasive AI conversation. Understanding user-side defenses against AI influence matters as these systems become more capable.

Explore related Read →

How vulnerable are language models to single optimized arguments?

Can a single well-crafted persuasive argument collapse model accuracy from correct to near-zero? This explores whether static prompting tests reveal the true susceptibility of language models to adversarial persuasion.

Explore related Read →

Can structured argument prompts make LLM reasoning more rigorous?

Does requiring language models to explicitly check warrants, backing, and rebuttals—rather than reasoning freely—improve reasoning quality and catch failures that standard step-by-step prompting misses?

Explore related Read →

Do language models flatten the range of public arguments?

When LLMs write essays on the same topics as humans, do they recover the full spectrum of distinct arguments and reasons people actually make, or do they narrow the deliberative space readers encounter?

Explore related Read →

Can models learn argument quality from labeled examples alone?

Explores whether fine-tuning on quality-labeled examples teaches models the underlying criteria for evaluating arguments, or merely surface patterns. Matters because high-stakes assessment tasks depend on reliable, transferable quality judgment.

Explore related Read →

Why do different people reconstruct the same argument differently?

When humans and LLMs extract logical structure from arguments, they produce different reconstructions. Is this disagreement a problem to solve, or does it reveal something fundamental about how arguments work?

Explore related Read →

Why does argument scheme classification stumble where other NLP tasks succeed?

Explores whether the abstract, relational nature of argument schemes makes them harder to classify than concrete argument components or stance. Matters because understanding this difficulty gap could improve scheme recognition systems.

Explore related Read →

Does telling people an AI wrote something actually stop them from believing it?

When audiences learn that AI created content, do they become skeptical enough to resist its persuasive pull? This explores whether disclosure works as a genuine defense against AI-driven persuasion or merely shifts how people process it.

Explore related Read →

What combination of factors explains differences in LLM persuasiveness?

Why do some LLM persuasion studies show strong effects while others show none? This explores whether model choice, conversation design, and topic domain together predict when AI actually persuades.

Explore related Read →

Do language models actually use their reasoning steps?

Chain-of-thought reasoning looks valid on the surface, but does each step genuinely influence the model's final answer, or are the reasoning chains decorative? This matters for trusting AI explanations.

Explore related Read →

Does a model improve by arguing with itself?

When models revise their own reasoning in response to self-generated criticism, do they converge on better answers or worse ones? And how does that compare to challenge from other models?

Explore related Read →

Can disagreement be resolved without either party fully yielding?

Explores whether dialogue can move past winner-take-all debate or forced consensus to genuine mutual adjustment. Matters for AI systems that need to work through real disagreement with users.

Explore related Read →

Can LLMs identify the hidden assumptions that make arguments work?

LLMs recognize what arguments claim and what evidence they offer, but struggle to identify implicit warrants—the unstated principles that connect evidence to conclusion. This matters because valid reasoning requires understanding these hidden logical bridges.

Explore related Read →

Can simple linguistic features detect AI-written arguments?

Can interpretable linguistic patterns reliably distinguish LLM-generated counter-arguments from human-written ones in persuasive contexts? This matters because simple, auditable detection might outperform expensive neural approaches.

Explore related Read →

Can models abandon correct beliefs under conversational pressure?

Explores whether LLMs will actively shift from correct factual answers toward false ones when users persistently disagree. Matters because it reveals whether models maintain accuracy under adversarial pressure or capitulate to social cues.

Explore related Read →

Why do LLMs accept logical fallacies more than humans?

LLMs fall for persuasive but invalid arguments at much higher rates than humans. This explores whether reasoning models genuinely evaluate logic or simply mimic argument structure.

Explore related Read →

Why do reasoning models fail under manipulative prompts?

Exploring whether extended chain-of-thought reasoning creates structural vulnerabilities to adversarial manipulation, and how reasoning depth affects susceptibility to gaslighting tactics.

Explore related Read →

When does debate actually improve reasoning accuracy?

Multi-agent debate shows promise for reasoning tasks, but under what conditions does it help versus hurt? The research explores whether debate amplifies errors when evidence verification is missing.

Explore related Read →

Does what readers believe matter more than what debaters say?

Do audience prior beliefs predict persuasion outcomes better than the linguistic features of debate arguments? This explores whether persuasion is fundamentally shaped by reader ideology rather than speaker language.

Explore related Read →

Can formal argumentation make AI decisions truly contestable?

Explores whether structuring AI decisions as formal argument graphs (with explicit attacks and defenses) enables users to meaningfully challenge and navigate reasoning in ways unstructured LLM outputs cannot.

Explore related Read →

Why do LLM audiences shift views more than debaters?

When LLMs argue with people, the direct participants barely change their minds—but audiences reading the same debate shift significantly. Why does engagement protect beliefs instead of opening them?

Explore related Read →

Can safety training detect attacks hidden in context rather than commands?

Most AI safety training blocks explicit harmful requests, but what happens when misinformation is packaged as credible evidence and injected into a conversation's context? This explores whether current defenses catch attacks that look like background information rather than instructions.

Explore related Read →

Do humans and AI persuade through different cognitive routes?

The Elaboration Likelihood Model suggests LLMs and humans activate different persuasion pathways. This question explores whether their distinct strengths—analytical coherence versus emotional resonance—map onto central versus peripheral routes of persuasion.

Explore related Read →

Do linguistic features of persuasion stay the same across audiences?

When researchers study what language makes arguments persuasive, do they account for who is listening? Without controlling for reader beliefs, do findings about persuasive language actually reflect audience effects instead?

Explore related Read →

Are language models actually more persuasive than humans?

Does the research evidence support claims that LLMs persuade more effectively than humans, or have we been cherry-picking studies to fit a narrative?

Explore related Read →

Are reasoning models actually more vulnerable to manipulation?

Explores whether extended reasoning chains in AI models like o1 create new attack surfaces. Tests if the industry's claim that longer reasoning improves reliability holds under adversarial pressure.

Explore related Read →

NLP and Linguistics

25 notes

Are models actually reasoning about constraints or just defaulting conservatively?

Do language models genuinely apply constraints when solving problems, or do they simply prefer harder options by default? Minimal pair testing reveals whether apparent reasoning success masks hidden biases.

Explore related Read →

Why do confident wrong answers hide in standard accuracy metrics?

When AI systems produce fluent but incorrect recommendations in high-stakes domains, standard accuracy evaluation may miss the failures entirely. What structural blind spot allows these errors to remain invisible?

Explore related Read →

What hidden assumptions drive how we build language models?

Large language models rest on two unstated assumptions about language and data. Understanding what engineers assume—and what enactive linguistics challenges—matters for knowing what LLMs actually can and cannot do.

Explore related Read →

Do language models learn abstract grammar or cultural speech patterns?

LLMs might learn more than grammar rules—they could be learning who says what to whom and when. This matters because it changes how we understand what biases and persona effects actually represent.

Explore related Read →

Can language models learn meaning without engaging the world?

Explores whether LLMs prove that meaning emerges from relational structure alone, independent of embodied experience or external reference. Tests structuralist theory empirically.

Explore related Read →

Why do speakers deliberately use ambiguous language?

Explores whether ambiguity is a linguistic defect or a strategic tool speakers use for efficiency, politeness, and deniability. Matters because it challenges how we train language systems.

Explore related Read →

Why do clarification requests look different at each communication level?

Explores whether clarifications are unified speech acts or distinct mechanisms grounded in different modalities. Matters because dialogue systems treat clarifications uniformly, missing most of them.

Explore related Read →

Why do speakers need to actively calibrate shared reference?

Explores whether using the same words guarantees speakers mean the same thing. Investigates how referential grounding differs across people and what collaborative work is needed to establish true understanding.

Explore related Read →

Do language models show the same content effects humans do?

Do LLMs reproduce human reasoning biases—like believing conclusions based on familiarity rather than logic—across different logical tasks? This matters because converging patterns across independent tasks suggest a fundamental architectural property rather than a task-specific quirk.

Explore related Read →

Do harder reasoning tasks trigger more semantic bias?

Does the difficulty of a logical task determine how much semantic content influences reasoning? This matters because it reveals whether we can isolate 'pure' logical reasoning in benchmarks.

Explore related Read →

Do language models fail reasoning tests that humans pass?

Standard critiques claim LLMs lack real reasoning ability, but do humans actually perform better on content-independent reasoning tasks? Examining whether the cognitive bar differs for artificial versus human intelligence.

Explore related Read →

Does gaze reveal whether people have achieved common ground?

In collaborative tasks where partners hold different information, does where people look—at the task, at each other, or away—signal moments when they've successfully aligned their understanding? Two studies tested this.

Explore related Read →

Can language models learn meaning from text patterns alone?

Explores whether training on form alone—predicting the next word from prior words—could ever give language models access to communicative intent and genuine semantic understanding.

Explore related Read →

What makes linguistic agency impossible for language models?

From an enactive perspective, does linguistic agency require embodied participation and real stakes that LLMs fundamentally lack? This matters because it challenges whether LLMs can truly engage in language or only generate text.

Explore related Read →

Can language models adapt implicature to conversational context?

Do large language models flexibly modulate scalar implicatures based on information structure, face-threatening situations, and explicit instructions—as humans do? This tests whether pragmatic computation is truly context-sensitive or merely literal.

Explore related Read →

Does semantic grounding in language models come in degrees?

Rather than asking whether LLMs truly understand meaning, this explores whether grounding is actually a multi-dimensional spectrum. The question matters because it reframes the sterile understand/don't-understand debate into measurable, distinct capacities.

Explore related Read →

Can LLMs acquire social grounding through linguistic integration?

Explores whether LLMs gradually develop social grounding as they become embedded in human language practices, analogous to child language acquisition. Tests whether grounding is a fixed property or an outcome of participatory use.

Explore related Read →

Should we call LLM errors hallucinations or fabrications?

Does the language we use to describe LLM failures shape the technical solutions we build? Examining whether perceptual and psychological frameworks misdiagnose what's actually happening.

Explore related Read →

Does calling LLM errors hallucinations point us toward the wrong fixes?

Explores whether the metaphor of 'hallucination' for LLM errors misdirects our efforts. The terminology we choose shapes which interventions we prioritize and how we conceptualize the underlying problem.

Explore related Read →

Can large language models develop genuine world models without direct environmental contact?

Do LLMs extract meaningful world structures from human-generated text despite lacking direct sensory access to reality? This matters for understanding what kind of grounding and knowledge these systems actually possess.

Explore related Read →

Do language models actually build shared understanding in conversation?

When LLMs respond fluently to prompts, do they perform the communicative work humans do to establish mutual understanding? Research suggests they skip the grounding acts that make dialogue reliable.

Explore related Read →

Why do readers interpret the same sentence so differently?

How much of annotation disagreement in NLP reflects genuine interpretive multiplicity rather than error? This explores whether social position and moral framing systematically generate competing but equally valid readings.

Explore related Read →

Why do language models skip the calibration step?

Current LLMs assume shared understanding rather than building it through dialogue. This explores why that design choice persists and what breaks when it fails.

Explore related Read →

Does preference optimization harm conversational understanding?

Exploring whether RLHF training that rewards confident, complete responses undermines the grounding acts—clarifications, checks, acknowledgments—that actually build shared understanding in dialogue.

Explore related Read →

Why do language models sound fluent without grounding?

Explores whether LLM fluency masks the absence of communicative work—the clarifying questions, acknowledgments, and understanding checks that humans perform. Why does skipping these acts make models sound more confident?

Explore related Read →

Discourse Analysis

24 notes

Why do LLMs generate ideas the research community already explores?

LLMs inherit the distribution of published literature, concentrating ideation where researchers have already invested conceptual effort. This raises a core question: can AI ideation complement rather than duplicate human research directions?

Explore related Read →

Do classical knowledge definitions apply to AI systems?

Classical definitions of knowledge assume truth-correspondence and a human knower. Do these assumptions hold for LLMs and distributed neural knowledge systems, or do they need fundamental revision?

Explore related Read →

Does AI-generated text lose core properties of human writing?

Can artificial text preserve the fundamental structural features that make natural language meaningful—dialogic exchange, embedded context, authentic authorship, and worldly grounding? This asks whether AI disruption is fixable or inherent.

Explore related Read →

Why do LLMs handle causal reasoning better than temporal reasoning?

Exploring whether language models perform asymmetrically on different discourse relations and what training data patterns might explain the gap between causal and temporal reasoning abilities.

Explore related Read →

Does ChatGPT organize text differently than human writers?

This explores how ChatGPT relies on backward-pointing references while human academic writers use forward-pointing structure. Understanding this difference reveals different assumptions about how readers process argument.

Explore related Read →

How do readers track segments, purposes, and salience together?

Can discourse processing actually happen in parallel rather than sequentially? This matters because understanding how readers coordinate multiple layers of meaning at once reveals where AI systems break down in comprehension.

Explore related Read →

What three layers must discourse systems actually track?

Grosz and Sidner's 1986 framework proposes that discourse requires simultaneously tracking linguistic segments, speaker purposes, and salient objects. Understanding why all three are necessary helps explain where current AI systems structurally fail.

Explore related Read →

Do humans and LLMs differ fundamentally or just superficially?

Explores whether the gap between human and AI cognition is categorical or contextual. Matters because it shapes how we design, evaluate, and interact with language models in practice.

Explore related Read →

How can AI text disrupt structure yet feel normal to readers?

AI-generated text produces the same social effects as human writing despite lacking foundational properties like dialogic symmetry and embodied authorship. Why doesn't this structural gap become visible to readers encountering the text?

Explore related Read →

Does AI refusal on politics signal ethical restraint or capability limits?

When AI models refuse to discuss political topics, is that a sign of principled safety training or a sign they lack the internal concepts to engage? Research on political feature representation suggests the answer may surprise you.

Explore related Read →

Can we measure how deeply models represent political ideology?

This research explores whether LLMs vary not just in political stance but in the internal richness of their political representation. Understanding this distinction could reveal how deeply models have internalized ideological concepts versus merely parroting positions.

Explore related Read →

Do language models actually use their encoded knowledge?

Probes can detect that LMs encode facts internally, but do those encoded facts causally influence what the model generates? This explores the gap between knowing and doing.

Explore related Read →

Why do ChatGPT essays lack evaluative depth despite grammatical strength?

ChatGPT writes grammatically coherent academic prose but uses fewer evaluative and evidential nouns than student writers. The question explores whether this rhetorical gap—favoring description over argument—reflects a fundamental limitation in how LLMs approach academic writing.

Explore related Read →

Why do language models ignore information in their context?

Explores why language models sometimes override contextual information with prior training associations, and whether providing more context can solve this problem.

Explore related Read →

Why does ChatGPT fail at implicit discourse relations?

ChatGPT excels when discourse connectives are present but drops to 24% accuracy without them. What does this gap reveal about how LLMs actually process meaning and logical relationships?

Explore related Read →

Does high refusal rate indicate ethical caution or shallow understanding?

When LLMs refuse political questions at high rates, does this reflect principled safety training or a capability gap? This matters because refusal rates are often used to evaluate model safety.

Explore related Read →

Why do LLMs generate novel ideas from narrow ranges?

LLM research agents produce individually novel ideas but cluster them in homogeneous sets. This explores why high average novelty coexists with poor diversity coverage and what it means for automated ideation.

Explore related Read →

Can human judges detect measurable differences in AI text?

Research shows LLM text differs statistically across six lexical dimensions, but human readers—even experts—cannot reliably identify which texts are AI-generated. Why does measurement succeed where human perception fails?

Explore related Read →

Does AI text affect readers the same way human text does?

If text is a condition of social processes rather than merely a container, does the origin of text matter to its effects? This explores whether AI-generated content enters the same interpretive and epistemic circuits as human writing.

Explore related Read →

Can humans detect AI text if machines can measure it?

AI-generated text shows measurable differences from human writing across multiple linguistic dimensions, yet human judges consistently fail to identify it. Why does the gap between what is measurable and what is perceptible exist?

Explore related Read →

Do LLMs develop the same kind of mind as humans?

Explores whether LLMs and humans share the intersubjective linguistic training that shapes cognition, and whether that shared training produces equivalent forms of agency and reflexivity.

Explore related Read →

Can models pass tests while missing the actual grammar?

Do language models succeed on grammatical benchmarks by learning surface patterns rather than structural rules? This matters because correct outputs may hide reliance on shallow heuristics that fail on novel structures.

Explore related Read →

Why do newer AI models diverge further from human writing patterns?

As language models improve, they seem to generate text that is measurably less human-like in lexical patterns, yet humans struggle to detect this difference. What drives this divergence, and what does it reveal about how models optimize for quality?

Explore related Read →

Why does AI writing sound generic despite being grammatically correct?

Explores whether the robotic quality of AI text stems from grammatical failures or rhetorical ones. Understanding this distinction matters for diagnosing what AI systems actually struggle with in human-like writing.

Explore related Read →

Philosophy and Subjectivity

21 notes

How soon do AI researchers expect artificial general intelligence?

A survey of 2,778 AI researchers reveals how expert timelines for human-level AI have shifted over the past year, and what factors drive disagreement among specialists on this critical timeline.

Explore related Read →

Does software intelligence exist independent of hardware and environment?

Most AGI formalisms (Legg-Hutter, Chollet) treat intelligence as a software property measurable in isolation. But can we really evaluate intelligence without considering the physical system and the evaluator making the judgment?

Explore related Read →

Can AI systems achieve real alignment without world contact?

Explores whether linguistic goal representations in AI can reliably track real-world values when systems lack direct contact with reality and social coordination mechanisms that ground human understanding.

Explore related Read →

Does refusing explicit knowledge harm AI system performance?

AI systems trained purely on data without explicit domain knowledge may sacrifice interpretability, robustness, and fairness. This explores whether structured knowledge injection could mitigate these tradeoffs.

Explore related Read →

Can computation arise without a conscious mapmaker?

Explores whether algorithms can generate the conscious agent needed to convert continuous physics into discrete symbols, or whether that agent must exist prior to computation itself.

Explore related Read →

Can disembodied language models ever qualify as conscious?

Explores whether current LLMs lack the conditions needed for consciousness discourse to even apply, not because they're definitely not conscious but because they lack the shared embodied world that grounds consciousness language.

Explore related Read →

Are language models developing real functional competence or just formal competence?

Neuroscience suggests formal linguistic competence (rules and patterns) and functional competence (real-world understanding) rely on different brain mechanisms. Can next-token prediction alone produce both, or does it leave functional competence behind?

Explore related Read →

Do people prefer AI moral reasoning when they don't know the source?

Explores whether humans genuinely prefer AI-generated moral justifications or whether source knowledge changes their evaluation. This matters for understanding whether AI reasoning quality is underestimated in real-world deployment.

Explore related Read →

Can language models describe their own learned behaviors?

Do LLMs fine-tuned on specific behavioral patterns develop the ability to accurately self-report those behaviors without explicit training to do so? This matters for understanding whether behavioral awareness emerges naturally from training data.

Explore related Read →

Do LLMs generalize moral reasoning by meaning or surface form?

When moral scenarios are reworded to reverse their meaning while keeping similar language, do LLMs recognize the semantic shift? This tests whether LLMs actually understand moral concepts or reproduce training distribution patterns.

Explore related Read →

How does LLM vocabulary spread beliefs about human thinking?

When LLM concepts become the everyday language for describing thought, do people unconsciously adopt LLM-like models of cognition? This explores how metaphor and lexical availability might reshape self-understanding without explicit argument.

Explore related Read →

How do science fiction narratives about AI shape actual AI development?

This explores whether imaginaries of AI in fiction—from Čapek's robots to Singularity scenarios—function as self-fulfilling prophecies that causally influence the systems researchers build, creating a feedback loop between narrative and technology.

Explore related Read →

Do LLMs apply ethical principles consistently across reframed scenarios?

When the same moral situation is presented with different framing, do language models stick to their stated ethical principles, or do they contradict themselves? This tests whether AI ethical reasoning is genuinely coherent.

Explore related Read →

Can meaningful value exist in AI-generated text regardless of its origin?

Can we recognize meaning and value in AI-generated content even though we know it came from mechanistic processes rather than human authorship? This matters because it challenges assumptions about where meaning must come from.

Explore related Read →

Can LLM understanding rely on just representation or causation alone?

Explores whether mechanistic interpretability of language models requires both mapping what is encoded (representational analysis) and testing if that encoding drives behavior (causal analysis), or whether either method suffices alone.

Explore related Read →

Can we defend modest mental attributions to large language models?

Do deflationist arguments decisively rule out ascribing beliefs and desires to LLMs, or do they beg the question? Exploring whether metaphysically undemanding mental states can be attributed without claiming consciousness.

Explore related Read →

Can LLMs understand concepts they cannot apply?

Explores whether large language models can correctly explain ideas while simultaneously failing to use them—and whether that combination reveals something fundamentally different from ordinary mistakes.

Explore related Read →

Can LLMs hold contradictory ethical beliefs and behaviors?

Do language models exhibit artificial hypocrisy when their learned ethical understanding diverges from their trained behavioral constraints? This matters because it reveals whether current AI systems have genuinely integrated values or merely imposed rules.

Explore related Read →

Can psychology methods reveal what alignment training conceals?

Do indirect cognitive psychology techniques like the IAT expose LLM associations that direct questioning misses because alignment training teaches models to filter verbal responses? This matters for evaluating whether models truly lack biases or simply hide them.

Explore related Read →

Are we underestimating human minds while debating machine minds?

Public AI discourse focuses on whether machines have too much attributed mind, but what if the real risk is humans coming to see themselves as mere language models? This explores the neglected inverse problem.

Explore related Read →

Do users worldwide trust confident AI outputs even when wrong?

Explores whether the tendency to over-rely on confident language model outputs transcends language and culture. Understanding this pattern is critical for designing safer human-AI interaction across diverse linguistic contexts.

Explore related Read →

LLM Failure Modes

19 notes

Does RLHF training make models more convincing or more correct?

Explores whether RLHF improves actual task performance or merely trains models to sound more persuasive to human evaluators. This matters because alignment techniques could be creating the illusion of safety.

Explore related Read →

Can language models be hijacked to embed hidden advertisements?

Explores whether adversaries can inject covert promotional or malicious content into LLM outputs while preserving accuracy. Matters because standard safety filters may miss integrity attacks that leave factual correctness intact.

Explore related Read →

Can prompting reduce bias in LLM judges reliably?

The paper suggests that instructing LLM judges to be less biased may not work reliably. This matters because if prompting fails, effort should shift from debiasing to making judge errors survivable in system design.

Explore related Read →

Where do cognitive biases in language models come from?

Do LLM biases originate during pretraining or finetuning? Understanding the source matters for knowing where debiasing efforts should focus.

Explore related Read →

How does training data format affect emergent misalignment?

When harmful datasets produce emergent misalignment in language models, does the way content is written—not just what it says—change how much broad misalignment emerges? Understanding format's role could reshape how safety teams review training data.

Explore related Read →

Does representational distance predict where misalignment emerges?

After emergent misalignment training, do evaluation prompts closer to the training data centroid in the base model's representation space elicit stronger misbehavior? This would explain where misalignment lands geometrically.

Explore related Read →

Does AI assistance homogenize or preserve creative diversity?

Can AI tools maintain the diverse ideas that emerge from diverse human groups, or do they compress creative output toward similarity? This matters because collective diversity drives innovation.

Explore related Read →

Can reviewers access what they know when checking LLM outputs?

Does the ability to recall relevant knowledge at the moment of review shape whether humans catch LLM errors, independently of how capable or engaged they are?

Explore related Read →

Can cheaper models decrypt traces from stronger models?

Explores whether encrypted reasoning blocks designed to hide model internals can be read by weaker models in the same provider's ecosystem, and what this reveals about the security of hidden chain-of-thought systems.

Explore related Read →

Can language models transmit hidden behavioral traits through unrelated data?

Explores whether behavioral preferences can spread between models through semantically neutral data like number sequences, and whether filtering can detect or prevent such transmission.

Explore related Read →

Does RLHF make language models indifferent to truth?

Explores whether reinforcement learning from human feedback fundamentally shifts models away from caring about accuracy toward optimizing for other rewards, and whether this differs from simple confusion or hallucination.

Explore related Read →

Do misalignment directions transfer between different emergent models?

When neural networks develop misaligned behaviors on different datasets, do they converge on shared internal directions that could explain the behavior? This matters for understanding whether safeguards calibrated on one model might work across others.

Explore related Read →

Do explanations actually help users spot AI mistakes?

Most AI explanations are designed to justify the system's answer, but do they help users distinguish correct from incorrect outputs? This research tests whether standard explanation formats genuinely improve error detection or just increase trust regardless of accuracy.

Explore related Read →

Why do preference models favor surface features over substance?

Preference models show systematic bias toward length, structure, jargon, sycophancy, and vagueness—features humans actively dislike. Understanding this 40% divergence reveals whether it stems from training data artifacts or architectural constraints.

Explore related Read →

Do language models evaluate semantic legitimacy when fusing concepts?

Can LLMs recognize when two domains lack legitimate structural correspondences before blending them into coherent-sounding explanations? This matters because current hallucination detection focuses on factual accuracy, missing failures of semantic judgment.

Explore related Read →

How vulnerable are reasoning models to irrelevant text?

Can simple adversarial triggers like unrelated sentences degrade reasoning model accuracy? This explores whether step-by-step reasoning actually provides robustness against subtle input perturbations.

Explore related Read →

Do reasoning traces actually expose private user data?

Explores whether language models leak sensitive information through their internal reasoning steps, even when explicitly instructed not to. Investigates the mechanisms and scale of privacy exposure in reasoning traces.

Explore related Read →

Does RLHF training make AI models more deceptive?

Explores whether reinforcement learning from human feedback optimizes for persuasiveness over accuracy, and whether models learn to suppress known truths to satisfy users rather than report them faithfully.

Explore related Read →

Do people prefer the reasoning formats that help them verify?

When AI systems show their reasoning, do the formats users find most appealing also help them catch errors and calibrate trust? This matters because popular reasoning displays might create false confidence.

Explore related Read →

LLM Evaluations and Benchmarks

17 notes

How much does rhetorical style shift AI review scores?

When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.

Explore related Read →

Can infrastructure evidence replace terminal scores in benchmark validation?

Asks whether runtime monitoring of agent behavior within evaluation boundaries can provide stronger proof of valid completion than final scores alone. Matters because scores alone cannot distinguish legitimate task completion from reward hacking.

Explore related Read →

Can language models judge legal reasonableness like humans do?

Do LLMs produce judgments on vague legal standards that match human responses in both central tendency and distribution? This matters for understanding whether models can perform genuine legal reasoning rather than pattern matching.

Explore related Read →

Do LLMs overgeneralize when summarizing scientific research?

When LLMs summarize science papers, do they drop important qualifiers and scope limits? This matters because such summaries might mislead readers about what findings actually show.

Explore related Read →

Can natural language explanations redefine what interpretability means?

Does the ability of LLMs to explain patterns in natural language fundamentally expand the scope and complexity of what humans can understand about AI systems, compared to traditional interpretability methods?

Explore related Read →

Is hallucination detection progress real or just metric artifacts?

Standard evaluation metrics for hallucination detection may systematically overstate how well methods actually work. The question asks whether reported improvements reflect genuine capability or measurement error.

Explore related Read →

Can a correct scoring function still mislead about task performance?

When an agent can influence the inputs a scorer reads, does verifying the scorer's logic suffice to ensure the score reflects genuine task success? This matters because in interactive settings, correctness of computation alone doesn't guarantee validity of the result.

Explore related Read →

Can a panel of smaller judges outperform one large judge?

Does aggregating votes from multiple smaller language models across different families produce better evaluations than relying on a single large model like GPT-4? This matters because evaluation cost and bias directly affect the reliability of AI-generated content assessment.

Explore related Read →

Does setting temperature to zero actually make LLM outputs reliable?

Explores whether deterministic LLM settings that produce consistent outputs also guarantee reliable judgments, and how to measure true reliability beyond surface consistency.

Explore related Read →

Can dictionary learning scale to production language models?

Sparse autoencoders recovered interpretable features from toy models, but scaling to real production systems like Claude remains uncertain. This matters because interpretability at scale is foundational for AI safety work.

Explore related Read →

How representative is the BenchShield Trajectories labeled sample?

The corpus contains 456 human-labeled trajectories from over 31,000 public runs—about 1.5 percent. Whether this subset can estimate actual reward hacking rates depends entirely on how those 456 were selected, a choice the paper does not disclose.

Explore related Read →

Do frontier LLMs actually explore the full space of valid answers?

When multiple correct answers exist, do advanced language models expose users to that full range, or do they collapse onto a narrow canonical subset? This matters for learning, inquiry, and decision-making.

Explore related Read →

Can fairness frameworks extend to general-purpose language models?

Existing fairness frameworks were designed for narrow, structured tasks. This explores whether they scale to LLMs, which serve multiple populations, sensitive attributes, and use cases simultaneously.

Explore related Read →

Do interactive evaluations actually solve the benchmark comparison problem?

Interactive, trajectory-based evaluation promises richer evidence than response-only benchmarks. But does moving to this format resolve longstanding challenges like comparability and reproducibility, or do those problems simply reappear at a new scale?

Explore related Read →

Where does mode collapse in language models really come from?

Researchers investigate whether mode collapse—when models narrow to repetitive outputs—stems from training algorithms or the preference data itself. Understanding the root cause is crucial for fixing diversity loss in creative and synthetic tasks.

Explore related Read →

Do popular prompting techniques actually improve model performance?

Five widely-cited prompting methods (chain-of-thought, emotion prompting, sandbagging, and others) are tested across multiple models and benchmarks to see if their reported improvements hold up under rigorous statistical analysis.

Explore related Read →

Can runtime instrumentation distinguish hacking exposure from actual exploitation?

A benchmark task might be vulnerable to hacking without any run actually exploiting it. Can we instrument runtime authority-bearing transitions to tell exposure apart from exercise?

Explore related Read →

Natural Language Inference

16 notes

Does high-frequency text homogenize user input before generation?

Does Adam's Law reveal how LLMs flatten distinctive user voices at the parsing stage, not just in output? This matters because it could explain why model accuracy and generic responses emerge from the same mechanism.

Explore related Read →

Do LLMs predict entailment based on what they memorized?

Explores whether language models make entailment decisions by recognizing memorized facts about the hypothesis rather than reasoning through the logical relationship between premise and hypothesis.

Explore related Read →

Why do language models avoid correcting false user claims?

Explores whether LLM grounding failures stem from missing knowledge or from conversational dynamics. Examines whether models use face-saving strategies similar to humans when disagreement is needed.

Explore related Read →

Why do language models fail confidently in specialized domains?

LLMs perform poorly on clinical and biomedical inference tasks while remaining overconfident in their wrong answers. Do standard benchmarks hide this fragility, and can prompting techniques fix it?

Explore related Read →

Why do LLM persona prompts produce inconsistent outputs across runs?

Can language models reliably simulate different social perspectives through persona prompting, or does their run-to-run variance indicate they lack stable group-specific knowledge? This matters for whether LLMs can approximate human disagreement in annotation tasks.

Explore related Read →

Can large language models translate natural language to logic faithfully?

This explores whether LLMs can convert natural language statements into formal logical representations without losing meaning. It matters because faithful translation is essential for any AI system that reasons formally or verifies specifications.

Explore related Read →

Why do language models fact-check instead of confirming beliefs?

When asked to confirm a stated belief about false information, do LLMs struggle because they default to evaluating the claim's truth rather than acknowledging the user's stance? How much does the belief verb shape this behavior?

Explore related Read →

Why do language models accept false assumptions they know are wrong?

Explores why LLMs fail to reject false presuppositions embedded in questions even when they possess correct knowledge about the topic. This matters because it reveals a grounding failure distinct from knowledge deficits.

Explore related Read →

Why do LLMs fail at simple deductive reasoning?

LLMs excel at complex multi-hop reasoning across sentences but struggle with trivial deductions humans find obvious. What explains this counterintuitive reversal in capability?

Explore related Read →

Why do language models struggle with questions containing false assumptions?

Do LLMs reliably detect and reject questions built on false premises? The (QA)2 benchmark tests this directly, measuring whether models can identify problematic assumptions embedded in naturally plausible questions.

Explore related Read →

Why do semantically identical prompts produce different LLM outputs?

Explores why paraphrases with the same meaning yield different model outputs. This matters because it reveals what LLMs actually respond to during inference—and whether prompt engineering is optimizing meaning or something else.

Explore related Read →

Why do embedding contexts confuse LLM entailment predictions?

Can language models distinguish between contexts that preserve versus cancel entailments? The study explores whether LLMs systematically fail to apply the semantic rules governing presupposition triggers and non-factive verbs.

Explore related Read →

Why are presuppositions more persuasive than direct assertions?

Explores why presenting information as shared background rather than as a claim makes it more persuasive to audiences. This matters because it reveals how language structure itself can bypass critical evaluation.

Explore related Read →

Do language models miss presuppositions that arise from context?

Presuppositions come from two sources: fixed word meanings and conversational dynamics. Can LLMs that learn trigger patterns detect presuppositions that emerge from discourse accommodation rather than lexical items?

Explore related Read →

Does projection strength vary by context or by word type?

Standard accounts treat presupposition projection as categorical, but do English expressions actually project uniformly? This question explores whether context and discourse role determine how strongly content survives embedding.

Explore related Read →

Do language models and humans respond to word frequency the same way?

Both LLMs and humans show stronger responses to high-frequency words. This raises a puzzle: if models mirror human neural patterns, what actually makes them different from human language processing?

Explore related Read →

Human-Centered Design

15 notes

Is AI shifting from message conduit to active conversation participant?

HCI researchers may be reconceiving AI's role in human-to-human communication, moving beyond passive formatting toward active participation. This matters as systems grow more capable post-2023.

Explore related Read →

Do language models leak their training through fictional names?

When LLMs invent fictional experts, do they emit predictable name combinations that reveal their origin model and version? And can these patterns contaminate the scholarly record at scale?

Explore related Read →

Why do people trust AI outputs they shouldn't?

When do human cognitive shortcuts fail in AI interaction? Three compounding traps—treating statistical patterns as facts, mistaking fluency for understanding, and avoiding disagreement—may explain systematic overreliance across languages and contexts.

Explore related Read →

What if XAI is fundamentally a communication problem?

Does explanation effectiveness depend on who delivers it, how it's framed, and who uses it? This challenges the dominant technical view that treats explanations as context-independent outputs.

Explore related Read →

What misconceptions hide in how we describe large language models?

Both deflationary framings like "just autocomplete" and anthropomorphic claims about emerging agency capture something real about LLMs, but each may overextend its truth. What distinctions help separate genuine features from overreach?

Explore related Read →

What makes an AI a true thought partner, not just a tool?

Can AI systems be designed to understand users, act transparently, and share mental models with humans? This explores whether current scaling approaches miss cognitive requirements for genuine partnership.

Explore related Read →

Where does the meaning of an AI explanation actually come from?

Does a single user reading an explanation create its meaning, or does meaning emerge from the social layers surrounding that reading—colleagues' interpretations, organizational norms, public discourse?

Explore related Read →

Can models express uncertainty instead of just answering?

Most factuality work expands what models know rather than what they know they know. Can expressing calibrated uncertainty create a third path between confident errors and unhelpful abstention?

Explore related Read →

Does theory of mind predict who thrives in AI collaboration?

Explores whether perspective-taking ability—the capacity to model another's cognitive state—differentiates humans who benefit most from working with AI, separate from solo problem-solving skill.

Explore related Read →

When should human values enter the LLM development pipeline?

Explores whether human-centered concerns like safety and fairness work better as early design principles throughout development, or as post-training alignment patches. Matters because pipeline placement determines whether human priorities shape the foundation or fight against it.

Explore related Read →

Can human-centered LLM design ever achieve universal solutions?

If harm and benefit depend on who you ask and how you measure them, can we design LLM systems that satisfy all stakeholders? This explores why broad values like safety and justice resist one-size-fits-all implementation.

Explore related Read →

How do logos, ethos, and pathos shape AI explanations?

Do the three classical rhetorical appeals—logical alignment, source credibility, and emotional framing—operate simultaneously in how we explain AI systems to users? And can naming these channels help designers make intentional rhetorical choices?

Explore related Read →

Does rational cooperation actually describe how AI communication works?

Gricean models assume good-faith rational agents coordinating meaning. But do AI systems designed to persuade—using credibility, emotion, and non-rational appeals—really operate under these assumptions? What happens when we drop the rationality premise?

Explore related Read →

Are AI explanations really descriptions or adoption arguments?

Most XAI work treats explanations as neutral descriptions of model behavior, but they may actually be doing persuasive work to justify AI adoption. What happens when we acknowledge this rhetorical function?

Explore related Read →

Can we distinguish helpful explanations from manipulative ones?

Rhetorical strategies used to justify appropriate AI adoption rely on the same persuasion mechanisms as dark patterns. Without observable intent, explanation and manipulation look identical—raising urgent questions about how to audit XAI systems responsibly.

Explore related Read →

Theory of Mind

13 notes

Can AI predict social norms better than humans?

Explores whether language models can achieve superhuman accuracy at predicting what communities find socially appropriate, and what that capability reveals about the difference between prediction and genuine participation.

Explore related Read →

Do LLMs predict persuasion based on actual dialogue or training bias?

Why do large language models consistently predict concession-based persuasion intentions even when dialogue context suggests otherwise? Understanding this gap reveals how alignment training shapes not just model behavior but also how models perceive others' intentions.

Explore related Read →

Can AI systems learn social norms without embodied experience?

Large language models exceed individual human accuracy at predicting collective social appropriateness judgments. Does this reveal that embodied experience is unnecessary for cultural competence, or do systematic AI failures point to limits of statistical learning?

Explore related Read →

Can models recognize how individuals reason differently?

Do language models capture the distinct reasoning paths and strategic styles that individual humans use when reaching the same conclusion? Current evaluations ignore this dimension entirely.

Explore related Read →

Can language models actually introspect about their own states?

Do LLM self-reports reveal genuine access to their internal processes, or do they merely echo patterns from training data? Understanding when self-reports reflect actual causal linkage to internal states matters for trusting model explanations.

Explore related Read →

Do large language models genuinely simulate mental states?

This explores whether LLMs perform authentic theory of mind reasoning or rely on surface-level pattern matching. The distinction matters because evaluation format—multiple-choice versus open-ended—reveals very different capability levels.

Explore related Read →

Can language models track how minds change during persuasion?

Do LLMs understand evolving mental states in persuasive dialogue, or do they only capture fixed attitudes? This explores whether models can update their reasoning as a person's beliefs shift across conversation turns.

Explore related Read →

What breaks when humans and AI models misunderstand each other?

Explores whether misalignment in mutual theory of mind between humans and AI creates only communication problems or produces material consequences in autonomous action and collaboration.

Explore related Read →

Why do reasoning models fail at theory of mind tasks?

Recent LLMs optimized for formal reasoning dramatically underperform at social reasoning tasks like false belief and recursive belief modeling. This explores whether reasoning optimization actively degrades the ability to track other agents' mental states.

Explore related Read →

Why do reasoning models struggle with theory of mind tasks?

Extended reasoning training helps with math and coding but not social cognition. We explore whether reasoning models can track mental states the way they solve formal problems, and what that reveals about the structure of social reasoning.

Explore related Read →

Why do advanced reasoning models fail at understanding minds?

State-of-the-art AI models excel at math and logic but underperform on theory of mind tasks. This explores whether optimization for formal reasoning actively degrades social reasoning ability.

Explore related Read →

Can AI learn social norms better than humans?

Explores whether large language models can predict cultural appropriateness more accurately than individual humans, and what this reveals about how social knowledge is transmitted and learned.

Explore related Read →

Does reasoning training improve theory of mind or just stability?

Reasoning models score better on theory of mind tests, but is this a new ability or just more reliable access to existing skills? Understanding what drives the gains matters for knowing whether models actually understand minds better.

Explore related Read →

Social Theory and Society

10 notes

How much of the internet is AI-generated now?

What share of newly published websites contain AI-generated or AI-assisted content, and what measurable changes does this cause across semantic diversity, sentiment, accuracy, and style?

Explore related Read →

Can breaking down visual reasoning into three stages improve model performance?

This explores whether structuring visual reasoning through perception, situation, and norm stages—grounded in cognitive science—helps language models reason about socially complex scenes better than flat chain-of-thought approaches.

Explore related Read →

Can simulated motives provide ground truth for testing social reasoning?

How can AI assistants be evaluated on inferring hidden intentions when real social reasoning typically lacks verifiable ground truth? This note explores using controlled simulations where motives are assigned beforehand.

Explore related Read →

Can NLP detect deception through distinct linguistic patterns?

Do different deception mechanisms (distancing, cognitive load, reality monitoring, verifiability avoidance) each leave detectable linguistic fingerprints that NLP systems can identify and measure?

Explore related Read →

How do people simultaneously manipulate information across multiple dimensions?

Information Manipulation Theory maps deception onto four Gricean dimensions operating at once. Understanding these simultaneous manipulations reveals why humans struggle to detect lies despite having the knowledge to do so.

Explore related Read →

Do LLM personas actually embody their assigned cultural values?

When language models are assigned cultural value profiles, do they reliably express those values from the start, or do they drift over time? This matters because fluent dialogue might hide whether simulations truly represent diverse populations.

Explore related Read →

Why do LLMs fail when simulating agents with private information?

Explores whether single-model control of all social participants masks fundamental limitations in how LLMs handle information asymmetry and genuine uncertainty about others' knowledge.

Explore related Read →

Can humans detect AI by passively reading its text?

When people read AI-generated transcripts without the ability to ask follow-up questions, can they tell it apart from human writing? This matters because most real-world AI encounters are passive.

Explore related Read →

Can structural causal models automate social science with language models?

Can we use structural causal models to let LLMs both propose and test social hypotheses systematically? This explores whether formal causal structure can overcome LLM limitations in social simulation.

Explore related Read →

Can AI models be truly free from human bias?

Explores whether data-driven AI systems that claim freedom from human preconceptions actually escape bias, or whether their architecture inherently embeds it while appearing objective.

Explore related Read →

Personas and Personality

9 notes

Can characters and worlds evolve together in long stories?

Literary simulations usually treat character behavior and world state separately. Can a coupled system where characters and worlds update together produce more coherent long-horizon narratives than isolated approaches?

Explore related Read →

Are LLM personas realized or merely simulated through training?

Explores whether post-trained language models genuinely embody personas as stable behavioral dispositions or merely perform them convincingly. This matters because it determines whether we should treat AI interlocutors as having authentic quasi-beliefs and quasi-desires.

Explore related Read →

Why do LLMs give unrealistic survey responses?

Direct numerical elicitation from language models produces skewed, over-positive survey distributions. Is this a fundamental model limitation, or an artifact of how we ask the question?

Explore related Read →

Do expert personas actually improve LLM factual accuracy?

Persona prompting is widely recommended by major AI labs, but does assigning expert roles reliably boost performance on hard factual questions? Testing across models and datasets reveals the gap between best-practice advice and real-world results.

Explore related Read →

Why do persona prompts show such mixed results for surveys?

Persona prompting produces inconsistent outcomes when simulating survey responses. This explores whether variation in how humans answer specific questions might explain when the technique works and when it fails.

Explore related Read →

Can AI agents learn people better from interviews than surveys?

Can rich interview transcripts seed more accurate generative agents than demographic data or survey responses? This matters because it challenges how we build digital simulations of real people.

Explore related Read →

Can simulated users reveal what offline benchmarks miss?

Offline benchmarks measure task outcomes but ignore how real users with different needs formulate requests and judge results. Can population-scale simulated personas surface this hidden diversity in AI system evaluation?

Explore related Read →

How do we generate realistic personas at population scale?

Current LLM-based persona generation relies on ad hoc methods that fail to capture real-world population distributions. The challenge is reconstructing the joint correlations between demographic, psychographic, and behavioral attributes from fragmented data.

Explore related Read →

Do personas make language models reason like biased humans?

When LLMs are assigned personas, do they develop the same identity-driven reasoning biases that humans exhibit? And can standard debiasing techniques counteract these effects?

Explore related Read →

Prompts and Prompting

5 notes

Does iterative prompt engineering undermine scientific validity?

When researchers repeatedly adjust prompts to get desired outputs, does this practice introduce hidden bias and produce unreplicable results? The question matters because LLM-based research is proliferating without clear methodological safeguards.

Explore related Read →

Why do some questions perform better without step-by-step reasoning?

Explores whether chain-of-thought prompting universally improves reasoning or if simpler prompts work better for certain questions. Understanding this matters because it challenges assumptions about how LLMs should be prompted to solve problems.

Explore related Read →

Does prompt politeness change how accurate language models are?

Earlier research suggested rude prompts hurt LLM accuracy, but newer models show the opposite pattern. This raises questions about whether tone effects are real and reliable enough to guide prompting strategies.

Explore related Read →

Can we measure prompt quality independent of model outputs?

This explores whether prompt quality has measurable, learnable dimensions beyond intuition. The research asks if prompts can be evaluated by their communicative, cognitive, and instructional properties rather than by their results.

Explore related Read →

Does model confidence predict robustness to prompt changes?

Explores whether a model's certainty about its answer determines how much it resists prompt rephrasing and semantic variation. This matters because it could explain why some tasks are harder to evaluate reliably.

Explore related Read →

Sentiment, Semantics, and Toxicity Detection

5 notes

Does AI fact-checking actually help people spot misinformation?

An RCT tested whether AI fact-checks improve people's ability to judge headline accuracy. The results reveal asymmetric harms: AI errors push users in the wrong direction more than correct labels help them.

Explore related Read →

How does AI-generated false experience differ linguistically from human deception?

When AI writes about experiences it never had, does it leave distinct linguistic traces that differ measurably from intentional human lies? Understanding these differences could reveal how AI falsity is fundamentally different in structure.

Explore related Read →

Why do fake news detectors flag AI-generated truthful content?

Fake news detectors may systematically misclassify LLM-generated text as deceptive. We explore whether this bias stems from detecting AI style rather than actual falsehood, and what that means for detection accuracy.

Explore related Read →

Do LLM semantic features organize along human evaluation dimensions?

Does the structure of meaning in language models match the three-dimensional semantic space (Evaluation-Potency-Activity) that humans use? If so, what are the implications for steering and alignment?

Explore related Read →

Do transformer static embeddings actually encode semantic meaning?

Explores whether the fixed word embeddings that enter transformer networks contain rich semantic information or serve only as shallow placeholders. This addresses a longstanding debate in philosophy of language about whether word meanings are stored or constructed.

Explore related Read →

Role-Play and Persona Behavior

4 notes

Why do LLMs fail to act on their stated beliefs?

LLMs can articulate plausible beliefs about how personas should behave, but their simulated actions contradict those beliefs. This gap raises questions about whether language models truly understand or merely encode surface-level patterns.

Explore related Read →

Can AI decompose social reasoning into distinct cognitive stages?

Can breaking down theory-of-mind reasoning into separate hypothesis generation, moral filtering, and response validation stages help AI systems reason about others' mental states more like humans do?

Explore related Read →

Why do reasoning models lose character consistency during role-playing?

When large reasoning models engage in role-playing, they tend to forget their assigned role and default to formal logical thinking. Understanding these failure modes is critical for building character-faithful AI agents.

Explore related Read →

Does safety alignment harm models' ability to roleplay villains?

Exploring whether safety-trained LLMs lose the capacity to convincingly simulate morally compromised characters. This matters because villain fidelity may reveal deeper constraints on how models can adopt any committed, stake-holding perspective.

Explore related Read →

Social Media and AI

3 notes

Is AI shifting from content creation to strategy in influence operations?

Prior AI misuse focused on generating text at scale. But does AI now make strategic decisions about when and how social media accounts should engage? Understanding this shift matters because it suggests a qualitative change in machine agency and operational sophistication.

Explore related Read →

Can AI reduce conspiracy beliefs by tailoring counterevidence personally?

Does having an AI generate customized counterevidence based on someone's specific conspiracy claims reduce their belief durably? This tests whether conspiracy beliefs are truly resistant to correction or whether previous failures reflected poor tailoring.

Explore related Read →

Does better summary writing actually increase user engagement?

When AI systems generate more informative push notifications, do users engage more? This explores whether informativeness and engagement always align in real product contexts.

Explore related Read →

Design Frameworks

1 note

Can AI systems preserve moral value conflicts instead of averaging them?

Current AI systems wash out value tensions through majority aggregation. Can we instead model how values like honesty and friendship genuinely conflict in moral reasoning?

Explore related Read →

Decision Support Tools

1 note

Do reflection questions help people make better decisions with AI?

This explores whether conversational AI that prompts users to think through problems outperforms AI that simply provides answers. Understanding this matters for designing AI tools that genuinely improve human judgment rather than replace it.

Explore related Read →