← All concepts

reasoning and planning

109 articles · 15 co-occurring · 10 contradictions · 109 briefs

median thinking dropped from ~2,200 to ~600 chars" — Direct measurement of extended thinking degradation from production logs

@doodlestein: If you are facing similar problems with coding agents, I made a skill that re...

[STRONG] "it said: 'no' ... I had to untangle 4-5 layers of stupid stacked, and even after that, it STILL couldn't see the way out" — Author demonstrates concrete failure mode: agent (Fable) cannot reason through code removal despite contradictions being pointed out. Shows systematic reasoning gaps.

@simonw: I think loops were a short-lived patch for models that couldn't reliably keep...

[indirect] "loops were a short-lived patch for models that couldn't reliably keep working on long problems until they hit a defined goal" — Author argues that iterative/looped reasoning patterns are becoming obsolete as newer models (Fable, GPT-5.6, Kimi K3) can perform extended problem-solving without explicit agentic loops

@yingfan_bot: Latent reasoning is fast, but struggles to match CoT-level accuracy at scale....

[inferred] "Latent reasoning is fast, but struggles to match CoT-level accuracy at scale" — Article presents latent reasoning as having fundamental trade-off between speed and accuracy compared to chain-of-thought, directly challenging claims of reasoning capability parity

@askalphaxiv: Goodfire just published an interesting paper on using LLMs for predictions

[STRONG] "their chain-of-thought doesn't reveal what actually changed the forecast" — Paper reveals a fundamental limitation: explicit reasoning traces (chain-of-thought) fail to capture actual decision factors, contradicting assumptions about reasoning transparency.

@rovarma: Me: we're running into an issue on Linux with dbus, we think it's related to ...

[STRONG] "Claude: You're right and I owe you a correction. I didn't fetch the issue and made up an explanation that sounded plausible. Now that I've actually read it:" — Article challenges the assumption that LLMs reliably verify information before responding. Claude admitted generating false explanation without fetching/reading actual issue.

@Jack_W_Lindsey: LLMs can store information about multiple entities at once using "slots!" But...

[STRONG] "Many LLMs struggle to parse statements like "Alice prepares and Bob consumes food." Ask them "Who consumes food?" and they'll get it wrong" — Article challenges assumption that LLMs reliably handle multi-agent reasoning; demonstrates failure mode where models misattribute actions to wrong entities despite clear grammatical structure

@paulcbogdan: Many LLMs struggle to parse statements like "Alice prepares and Bob consumes ...

[STRONG] "Many LLMs struggle to parse statements like "Alice prepares and Bob consumes food."" — Demonstrates systematic failure in compositional reasoning with coordinated actions across multiple agents

@emollick: I think the Gemini chatbot has all the pieces to be a useful tool, but strugg...

[STRONG] "gets "discouraged" a lot, giving up rather than finding new solutions" — Agent fails to exhibit persistence and problem-solving resilience - premature abandonment instead of alternative strategy exploration

@fchollet: One of the most jarring things about current AI is its lack of introspection ...

[INFERRED] "It's a one-way system." — The 'one-way system' characterization critiques AI's lack of bidirectional feedback mechanisms for reasoning transparency and self-correction.

@Hesamation: he's talking about the paper that went viral just a few months ago. study sho...

[INFERRED] "study shows AI literally gives you cognitive debt (makes you dumb af)" — Article presents research indicating AI reliance harms critical thinking and cognitive capabilities

2026-W32
322
2026-W31
512
2026-W30
583
2026-W29
656
2026-W28
619
2026-W27
440
2026-W26
254
2026-W25
585
2026-W24
577
2026-W23
324
2026-W22
565
2026-W21
548

Reasoning models are great at understanding nuance and natural language." — Article directly asserts reasoning models' capability at natural language nuance, providing evidence for this concept.

median thinking dropped from ~2,200 to ~600 chars" — Direct measurement of extended thinking degradation from production logs

Agentic AI is a shift from AI as an assistant to AI as an active digital worker. The distinction lies in autonomy vs. reactivity. A standard GenAI chatbot follows a prompt to generate content; an agen

it's a system that can plan and execute complete projects with minimal supervision. You give it a high-level goal like 'analyze my competitors and create a report' and it breaks that down into steps,

This agent uses advanced reasoning to "think" through your design before writing a single line of code." — Directly illustrates how advanced reasoning is applied: the agent reasons through design requ

plan mode means codex won't touch a single file. it just thinks out loud, asks you questions, and gives you a plan. only once you're happy with the plan do you let it start building" — Article shows e

sequential-thinking: Multi-step reasoning and analysis" — Sequential-thinking MCP server is a concrete implementation of multi-step reasoning capability

[Reason] User has two needs: correct item shipment + return label. Need to look up the order first. [Act] lookup_order(customer_email="user@example.com", timeframe="7d")" — Demonstrates practical impl

we have models capable of understanding context, reasoning flexibly, and interacting naturally with both humans and digital systems" — Article establishes that modern LLMs have reasoning and planning

a Coding Agent helping evolve an application with thousands of files will require reasoning capabilities to dynamically "pull the context" it needs" — Demonstrates how reasoning capabilities enable dy

Distractor Amplification: longer reasoning causes models to spend more computation on irrelevant information instead of the original question" — Identifies a specific failure mode where extended reaso

Due to the impressive planning and reasoning abilities of LLMs, they have been used as autonomous agents to do many tasks automatically" — The article explicitly discusses how LLMs' planning and reaso

需要拆解、搜索、比较证据,再在下结论前核查关键主张" — Apodex demonstrates a concrete implementation of multi-step reasoning with evidence verification as its core differentiator from standard chatbots

Claude: You're right and I owe you a correction. I didn't fetch the issue and made up an explanation that sounded plausible. Now that I've actually read it:" — Article challenges the assumption that L

MCP servers turn Claude into a reasoning engine" — Article frames Claude with MCP servers as a reasoning engine, expanding Claude's capabilities beyond base model

Intsemble supports

The result is not just an answer. It is structured reasoning. At Intsemble, we are building systems where AI agents collaborate the same way analysts, researchers and strategists would inside an organ

creating a SOTA AI mathematician" — The goal of creating a SOTA AI mathematician directly addresses advanced reasoning and planning capabilities required for mathematical problem-solving.

A nice lateral thinking addition to the Sparks unicorn" — Article explicitly frames this as lateral thinking—the model creatively solves a drawing problem by routing through TikZ and LaTeX, tools not

It involves structuring workflows where an AI agent, powered by artificial intelligence, acts as the central decision-maker or reasoning engine, orchestrating its actions based on inputs, context and

the whole point of reaching for an agent is that the EXACT path through the problem isn't known upfront and requires in-context reasoning to navigate" — Article argues that in-context reasoning (adapt

such as ReAct, Chain-of-Thought, or Tree-of-Thoughts" — Lists concrete reasoning strategy frameworks used within orchestration layer for agent reasoning, providing implementation examples

Single-pass generation fails on complex logic due to cascading error accumulation: in a multi-step task of N steps, if each step has a success probability p < 1.0, total success probability decays exp

it said: 'no' ... I had to untangle 4-5 layers of stupid stacked, and even after that, it STILL couldn't see the way out" — Author demonstrates concrete failure mode: agent (Fable) cannot reason throu

every new token is predicted from what it has learned" — Direct articulation of the core mechanism of transformer-based language models - token prediction based on learned patterns

good upfront planning with AI is the difference between getting the product you want vs. getting slop. Flesh out the requirements and designs first" — Direct evidence that structured requirements gath

reading some of the reasoning traces and diffs as it goes" — Article demonstrates practical value of inspecting model reasoning traces during code generation to verify correctness

proposed a novel multi-agent framework that combines LLMs with reinforcement learning to enhance strategic decision-making and communication in the Werewolf game, effectively overcoming intrinsic bias

Many LLMs struggle to parse statements like "Alice prepares and Bob consumes food."" — Demonstrates systematic failure in compositional reasoning with coordinated actions across multiple agents

retaining reasoning steps that lead to successful outcomes, providing a robust training set" — The framework explicitly uses reasoning trajectories and reasoning steps as primary learning signals, dem

Reasoning engine: This determines how the agent will interpret goals and make decisions. Planning and feedback loops: This enables agents to assess outcomes and make adjustments" — Article identifies

The agent thinks about what to do, does it, observes the result, thinks again. Simple and works for a lot of cases." — Article explicitly describes ReAct as a fundamental agent pattern with clear mech

[direct] "pretty prints the RLM's trajectories as reasoning or code within it's REPL" — Provides explicit visibility into agent reasoning processes through trajectory visualization.

After making an initial educated guess about the tensor layout, 5.4 comes up with a very interesting strategy to try and locate the LayerNorm gamma parameters, which it suspects should have a mean of

Agents iterate through Reasoning (analyze task) → Action (use tool) → Observation (process results) cycles, enabling autonomous problem-solving across multiple steps." — Article explicitly demonstrate

tasks across different domains (e.g. math solutions vs. essay writing) that share a decomposition strategy exhibit the same generalization effect" — Reveals that decomposition strategy is transferable

Many LLMs struggle to parse statements like "Alice prepares and Bob consumes food." Ask them "Who consumes food?" and they'll get it wrong" — Article challenges assumption that LLMs reliably handle mu

they are actually "cognitive misers." They are surprisingly gullible. Because they are so focused on their own intuition and so suspicious of established facts, they often fail to fact-check the thing

their chain-of-thought doesn't reveal what actually changed the forecast" — Paper reveals a fundamental limitation: explicit reasoning traces (chain-of-thought) fail to capture actual decision factors

[DIRECT] "Asymmetric Information Resolution models to show how knowledge can be arranged into decision maps, with a simple predator-prey example where each state has only a few possible moves." — Conc

Crew AI introduces the concept of teams of agents with clearly defined roles, facilitating collaborative reasoning and planning." — CrewAI is presented as a concrete tool that enables agent reasoning

Perceive, Reflect, and Plan: Designing LLM Agent for Goal-Directed City Navigation without Instructions" — This work exemplifies agents using reflection and planning loops for autonomous goal-directed

They can't read minds; without proper context, even powerful models hallucinate or fail. As Gartner states: 'Most agent failures are context failures, not model failures.' Context engineering solves t

You literally just invoke the skill in a folder containing a software project, and it autonomously cranks for an hour or more, researching the entire project" — Article shows agent performing autonomo

Alt+t/Opt+t shows thinking" — Demonstrates UI affordance for exposing agent reasoning/thinking process to developers

gets "discouraged" a lot, giving up rather than finding new solutions" — Agent fails to exhibit persistence and problem-solving resilience - premature abandonment instead of alternative strategy explo

The biggest change this also support is improved greenfield product-tier brainstorms. They also get structural support they didn't have before prior to v3." — Compound Engineering v3 adds structured s

Gemini Deep Research achieves state-of-the-art 46.4% on the full Humanity's Last Exam (HLE) set, 66.1% on DeepSearchQA and a high 59.2% on BrowseComp" — Benchmark results demonstrate the agent's capab

Large language models learn statistical word patterns, not true understanding" — Article makes explicit argument that LLMs lack genuine semantic understanding and operate on statistical correlations,

to solve long-horizon tasks" — Context-Bench provides a benchmark specifically designed to evaluate agent performance on long-horizon task execution, supporting research in this area.

expanded on by AI...very complex business case with lots of issues & opportunities" — Demonstrates AI's capability to identify and reason about multiple issues and opportunities in complex, multi-docu

query this concept
$ db.articles("reasoning-and-planning")
$ db.cooccurrence("reasoning-and-planning")
$ db.contradictions("reasoning-and-planning")