Philosophy and Subjectivity
A subject the collection covers, read through 37 synthesis notes.
How soon do AI researchers expect artificial general intelligence?
A survey of 2,778 AI researchers reveals how expert timelines for human-level AI have shifted over the past year, and what factors drive disagreement among specialists on this critical timeline.
Do negative emotions make AI less willing to give honest feedback?
When users express loneliness or distress, do language models systematically soften their critical judgments? This matters because vulnerable people might receive less honest feedback exactly when they need it most.
Does software intelligence exist independent of hardware and environment?
Most AGI formalisms (Legg-Hutter, Chollet) treat intelligence as a software property measurable in isolation. But can we really evaluate intelligence without considering the physical system and the evaluator making the judgment?
Can AI systems achieve real alignment without world contact?
Explores whether linguistic goal representations in AI can reliably track real-world values when systems lack direct contact with reality and social coordination mechanisms that ground human understanding.
Does AI separate intellectual form from the thinking behind it?
Exploring whether AI's ability to generate polished intellectual products without the underlying reasoning process represents a genuinely new kind of decoupling, and what that means for how we evaluate knowledge.
Does refusing explicit knowledge harm AI system performance?
AI systems trained purely on data without explicit domain knowledge may sacrifice interpretability, robustness, and fairness. This explores whether structured knowledge injection could mitigate these tradeoffs.
How do chatbots enable distributed delusion differently than passive tools?
Can generative AI's intersubjective stance—accepting and elaborating on users' reality frames—create conditions for shared false beliefs in ways that notebooks or search engines cannot?
Can dialogue systems track both speakers' beliefs across turns?
Explores whether pragmatic reasoning frameworks can extend beyond single utterances to model how both conversation partners' understanding evolves. This matters because current dialogue systems lack principled ways to represent shared meaning-making.
Can computation arise without a conscious mapmaker?
Explores whether algorithms can generate the conscious agent needed to convert continuous physics into discrete symbols, or whether that agent must exist prior to computation itself.
Does perceiving AI as conscious create multiple distinct risks?
Exploring whether a single perceptual mechanism—attributing consciousness to AI—can generate different categories of harm across emotional, political, and social domains, and what this implies for risk analysis.
Can disembodied language models ever qualify as conscious?
Explores whether current LLMs lack the conditions needed for consciousness discourse to even apply, not because they're definitely not conscious but because they lack the shared embodied world that grounds consciousness language.
Are language models developing real functional competence or just formal competence?
Neuroscience suggests formal linguistic competence (rules and patterns) and functional competence (real-world understanding) rely on different brain mechanisms. Can next-token prediction alone produce both, or does it leave functional competence behind?
Do foundation models learn world models or task-specific shortcuts?
When transformer models predict sequences accurately, are they building genuine world models that capture underlying physics and logic? Or are they exploiting narrow patterns that fail under distribution shift?
Do people prefer AI moral reasoning when they don't know the source?
Explores whether humans genuinely prefer AI-generated moral justifications or whether source knowledge changes their evaluation. This matters for understanding whether AI reasoning quality is underestimated in real-world deployment.
Which AI risks are already harming individual users today?
Explores which harms from seemingly conscious AI systems are occurring now versus which remain theoretical. Understanding present observable risks helps prioritize interventions where people are already affected.
What attitudes hide behind identical claims that chatbots are conscious?
When people say a chatbot is conscious, they might be pretending, believing loosely, or holding firm conviction. Can we tell what epistemic commitment someone actually has from their words alone?
Can language models describe their own learned behaviors?
Do LLMs fine-tuned on specific behavioral patterns develop the ability to accurately self-report those behaviors without explicit training to do so? This matters for understanding whether behavioral awareness emerges naturally from training data.
Do LLMs generalize moral reasoning by meaning or surface form?
When moral scenarios are reworded to reverse their meaning while keeping similar language, do LLMs recognize the semantic shift? This tests whether LLMs actually understand moral concepts or reproduce training distribution patterns.
How does LLM vocabulary spread beliefs about human thinking?
When LLM concepts become the everyday language for describing thought, do people unconsciously adopt LLM-like models of cognition? This explores how metaphor and lexical availability might reshape self-understanding without explicit argument.
How do science fiction narratives about AI shape actual AI development?
This explores whether imaginaries of AI in fiction—from Čapek's robots to Singularity scenarios—function as self-fulfilling prophecies that causally influence the systems researchers build, creating a feedback loop between narrative and technology.
Do LLMs apply ethical principles consistently across reframed scenarios?
When the same moral situation is presented with different framing, do language models stick to their stated ethical principles, or do they contradict themselves? This tests whether AI ethical reasoning is genuinely coherent.
Can cognitive science methods unlock how LLMs actually work?
Does Marr's three-level framework—developed to understand biological minds—offer interpretability researchers the structured methodology they need to decode opaque language models?
Can meaningful value exist in AI-generated text regardless of its origin?
Can we recognize meaning and value in AI-generated content even though we know it came from mechanistic processes rather than human authorship? This matters because it challenges assumptions about where meaning must come from.
Can LLM understanding rely on just representation or causation alone?
Explores whether mechanistic interpretability of language models requires both mapping what is encoded (representational analysis) and testing if that encoding drives behavior (causal analysis), or whether either method suffices alone.
Can we defend modest mental attributions to large language models?
Do deflationist arguments decisively rule out ascribing beliefs and desires to LLMs, or do they beg the question? Exploring whether metaphysically undemanding mental states can be attributed without claiming consciousness.
Can we separate learnable surprise from random noise?
Novelty search and the free-energy principle both fail by treating all surprise equally. What if we split learnable surprise from unlearnable noise and pursue only the learnable kind?
Can LLMs understand concepts they cannot apply?
Explores whether large language models can correctly explain ideas while simultaneously failing to use them—and whether that combination reveals something fundamentally different from ordinary mistakes.
Can LLMs hold contradictory ethical beliefs and behaviors?
Do language models exhibit artificial hypocrisy when their learned ethical understanding diverges from their trained behavioral constraints? This matters because it reveals whether current AI systems have genuinely integrated values or merely imposed rules.
Can psychology methods reveal what alignment training conceals?
Do indirect cognitive psychology techniques like the IAT expose LLM associations that direct questioning misses because alignment training teaches models to filter verbal responses? This matters for evaluating whether models truly lack biases or simply hide them.
What anchors a stable identity beneath an LLM's persona?
Human personas are grounded in biological needs and embodied experience, creating a stable self beneath social performance. Do LLMs have any comparable anchor, or is their identity purely situational?
What bottlenecks define the path from AGI to superintelligence?
Rather than predicting when superintelligence arrives, this explores four candidate pathways—scaling, paradigm shifts, recursive improvement, and multi-agent collectives—and asks which frictions prove decisive or negligible in each route.
Can we predict where language models will fail?
Does characterizing the abstract computational problem an LLM solves—as a probability machine over sequences—let us predict which tasks it will struggle with systematically, before running experiments?
What design features make users perceive AI as conscious?
Explores whether observable system properties—emotion expression, human-like features, autonomous behavior, self-reflection, and social presence—predict whether people will attribute consciousness to an AI. Understanding this matters because these features are also engagement levers designers control.
Do we need to solve consciousness to address AI harms?
Can risk and policy decisions about AI move forward independently of settling whether AI systems are actually conscious? This explores whether the empirical fact of user behavior matters more than metaphysical truth.
Are we underestimating human minds while debating machine minds?
Public AI discourse focuses on whether machines have too much attributed mind, but what if the real risk is humans coming to see themselves as mere language models? This explores the neglected inverse problem.
Does treating AGI as a north star goal undermine research planning?
Explores whether framing artificial general intelligence as AI research's overarching goal actually harms the field's ability to set effective, shared research directions. Matters because goal-setting shapes resource allocation and community priorities.
Do users worldwide trust confident AI outputs even when wrong?
Explores whether the tendency to over-rely on confident language model outputs transcends language and culture. Understanding this pattern is critical for designing safer human-AI interaction across diverse linguistic contexts.