AI in Practice

Computah: Make it a link

I spent a little time with the blog last night and pulled together two new site features using Claude Cowork. The last time I experimented significantly with Claude like this was to use Claude chat to build the link log from scratch, walking it through my thinking in plain language, then copying and pasting its suggestions into the backend of the site and hitting publish. This time I used Cowork, the tool that runs in the browser, and it clicked through the screens itself, fully taking on the execution of tasks. I have some coding skills, but not the kind these changes required. If I were taking this on, I would need YouTube, Hugo for Dummies, and my own personal IT guy, and still probably couldn’t pull it together.

Last night I asked a few things of Claude: I asked Claude Cowork to get into the backend of the micro.blog site and change my theme to link each line in the linklog back to my original post. The date field in the right column now links back to the original post for each link logged. No problem, easy request with easy execution. I made the request, confirmed the plan, and went about my business while Claude made the edits.

I also asked it to analyze my post content and suggest category tags for groups of content, then to label that content correctly in the backend. Claude reviewed about 500 posts and suggested I add a few new categories to my blog: AI in Practice, Books & Reading and Writing & Language. Then I set up some auto-filters to run at publish time to automatically categorize posts based on keywords moving forward.

Finally, I manually added the archive page, which lets you sort posts by category or year. This means I now have a functional archive here. Enjoy my anodyne thoughts, dear reader.

Adjacent to my day job, I’ve been toying with Claude Pro now for about a year; in my experience, it has improved significantly within the last six months. It’s not perfect: my requested edits were completed, but it also changed the CSS on the linklog so some of the text is too light to read, which I didn’t ask for and don’t want. But it has arguably extended my ability to execute on work that requires skills I don’t otherwise have (such as design and coding). What I do bring to the table is an expansive practical background in publishing and production, and all the language to describe it.

Gender, Power and AI: Wrestling for the soul of the network, again

Stanford’s Clayman Institute ran a virtual panel this morning called “Gender, Power, and Artificial Intelligence,” with Safiya Noble (UCLA), Catherine D’Ignazio (MIT), Angèle Christin (Stanford), and moderator Genevieve Smith, a Clayman Institute Postdoctoral Fellow. The panel applied principles from feminist tech studies to the current moment, and covered how gender norms get encoded in data and reproduced by AI systems, and discussed whether the technology has real capacity for equitable design and implementation at scale.

Noble’s argument throughout is that the governance conversation has gotten too high-level and universalizing while the actual outputs of these systems have profound day-to-day consequences for specific people today. She named the role of AI in the recent gerrymandering of Louisiana and Indiana as examples, and called for tripling down on long-term social science research about AI’s impacts. She also pointed out that philanthropy is retreating from feminist academic and organizational work because that work originates from the same dynamics that critique philanthropy itself, precisely at a point when this research is sorely needed. A lot of money is moving in AI, and very little of it is funding the people best positioned to study how it impacts everyone downstream.

D’Ignazio was asked directly whether feminist generative AI at scale is possible. Her answer was no, with caveats, given who owns the technology today and the current emphasis on profit motive. She suggested it is more important to consider how to organize around our relationship to technology, and how we might approach questions of profit and ownership, policy and decision-making, and data and tech governance.

She provided an example of a reasonable use case by walking us through a project from her Data + Feminism Lab. The example is documented at length in her recent book “Counting Feminicide: Data Feminism in Action,” where her team partnered with activists who scour news reports to document the gender-related killing of women and girls, including cisgender and transgender women. The lab built a very lightweight AI-based approach that streamlines the scanning and identification of news stories as possible cases to include in their project, supercharging their work (note: very similar to how the NYT uses AI to analyze data for reporting). In this example, the AI’s job is task-scoped, democratically co-determined with the people who use it, and small. Smith picked this up: there is an idea baked into the current LLM moment that AI must scale to make it marketable, and the alternative is using purpose-built models that are right-sized against a body of work.

Christin spoke at length about how embodiment is one of the primary focuses of feminist theory, and how AI perpetuates the “disembodied” illusion of technology, and how this dynamic shows up in everything from the marketing to UX to user comprehension. This spoke to my thoughts on how the single-interface design of LLM chat reproduces Haraway’s “god trick,” knowledge that presents as universal while concealing the specific and situated position it comes from.

The parallel I kept returning to, listening to this, is one I think about often with my own cohort of early bloggers, women who grew up alongside the rise of the internet — and then the rise of ad tech. The internet of the late 1990s and early 2000s was being shaped by several camps: writers, students, information architects, and user-centric researchers who saw it as an information access network and a space of possibility; entrepreneurs and opportunists who saw it as a channel for marketing, monetization and extraction; and a smaller boycott camp that wanted to limit and refuse the whole personal computing and digital revolution altogether.

It was generally considered weird to be a girl on a computer or a woman on the internet — so weird that many of our peers didn’t recognize us at all — and we were there anyway, making stuff, witnessing, learning, advocating, producing, influencing. So when I watch some of my old peers, many of whom are professional writers and academics today, treat LLMs as a question of refusal rather than a condition to engage with critically, I worry we are abdicating a responsibility at precisely the moment when our technical and rhetorical expertise applies. Their refusal has good logic: user-centric researchers and communities engaged extensively with the early internet and the extractive camp won anyway, so why expect a different outcome here?

But Noble’s work on algorithmic bias attributes that failure not to engagement, but to the institutional and financial disadvantages that user-centric approaches operated under relative to gargantuan commercial interests. David and Goliath. That gap does not close through abstention. Understanding the trade-offs around tech, producing knowledge and analysis that does not depend on investors and marketers to frame the platform and the questions, requires presence. Refusal cedes so much ground.

Overall, the recommendations from the panel were practical. Noble called for people with capital (and the political will to spend it) to consider how to put money toward socially responsible research and development. D’Ignazio called for alternative funding infrastructure outside of venture capital logic, and pointed at European digital sovereignty models as worthy of consideration here. She also gestured at the popular AI Skeptics reading group as one current example of mad-and-commiserating-as-organizing that is creating safe psychological space for people to talk about AI and its tradeoffs. Christin’s recommendation was community organizing, on the grounds that LLMs are unpopular with a lot of people who feel there is no space to say so, and that finding those spaces is itself worthy because it provides shared language and awareness of others’ knowledge and experiences.

Personally, it was refreshing to hear reflections on the work (and the feelings) of being inside institutions that are being reshaped by AI, and being responsible for some of how that reshaping gets communicated and absorbed. I’m thinking about the incredible value of interdisciplinary governance, and how the commitment to governance is a specific position, and all the margins to consider.

Further reading:

Catherine D’Ignazio and Lauren Klein, Data Feminism. The foundational text on applying intersectional feminist thinking to data science practice.

Catherine D’Ignazio, Counting Feminicide: Data Feminism in Action. Extended case study of the grassroots data activism project D’Ignazio described on the panel.

D’Ignazio et al., “Feminicide and Counterdata Production.” Research paper on the counterdata methodology behind the femicide tracking project.

D’Ignazio et al., “Data Feminism for AI.” Conference paper extending the data feminism framework to questions specific to AI systems.

Safiya Noble, Algorithms of Oppression. Noble’s study of how commercial search engines reinforce racism and sexism through their ranking systems.

Donna Haraway, “Situated Knowledges: The Science Question in Feminism and the Privilege of Partial Perspective” (1988). The original essay where Haraway introduces the god trick and the case for situated, embodied knowledge against the view from nowhere.

Reflections on teaching fiction writing in the age of AI, from a professor with ten years of classroom experience teaching writing at MIT.

A new study suggests that people who use AI for writing are more able to detect AI writing than automated scanner tools. My current LLM pet peeve is how they use language like load-bearing, structural and legible to describe most ideas.

Adventures in AI: I asked a Claude agent (new Opus, Pro plan) to build a Google Doc template with multiple tabs, using an existing doc as reference. It failed three times over two days, burned thru tokens, never worked with Drive. Eventually it spat out text for me to paste into a doc I made myself.

Fellow Madisonians, someone pulled together a website ranking local businesses in Madison by how local they are (by what criteria, idk). In my experience, this is one way we’re likely to see AI used in the next couple of years, via prototyping and/or executing ideas that result in dynamic websites.

Centaurs and Cyborgs on the Jagged Frontier by Ethan Mollick in 2023: “On some tasks AI is immensely powerful, and on others it fails completely or subtly. And, unless you use AI a lot, you won’t know which is which.”

Anecdotally, I’ve seen two family court cases where one party submitted full AI chats — prompts and colorful complaints included — as formal filings. The complaints wouldn’t pass muster with a real lawyer, but the conflict was nurtured by AI nonetheless. One was dinged for wasting the judge’s time.

I’ve posted a couple of times about instances I’m aware of where people are using AI in pro se court cases, especially family courts. A new study shows evidence of increasing numbers in pro se cases at the federal level, exacerbating existing bottlenecks. Many trade-offs abound here.

A professor asked students to self-report AI usage on their homework, leading to lots of confusion and uproar. Points aside, it’s clear people want more clarity up front about when and whether to use LLM tools. In the meantime, treating students like they’re guilty until proven innocent is a bad MO.

Timothy Chester offers some thoughts on the place of AI-assisted software development in a modern research university, and suggests that just because you can doesn’t necessarily mean you should.

I’m following a guy in TX who is using AI to write and illustrate children’s books whole cloth, then self-publishes using Amazon, and getting recognition in his region as a laudable children’s author. The books are categorically not good. It’s like people are rewarding his content strategy.

Silicon sampling is the practice of using LLMs to run surveys without talking to any people at all.

Innovations in scamming: Folks are predicting that AI will supercharge scams alongside any technical and administrative innovation. Here’s one example of an unethical use of AI, where an internet-based GLP-1 hub used AI to generate fake product images and before and after photos of smiling patients.

Testing a new feature I created using a mix of open source code and Claude, hoping I didn’t break my own site. I pulled together a dynamic link library using a Hugo partial and some shortcode that automatically catalogs all of my outbound links into sortable lists.

A screenshot of a Link library webpage displays a list of four links along with their titles and dates, sorted by Newest first under the category Higher Ed.A webpage lists blog-related links in a library format, sorted by newest first with dates.A webpage titled Link library displays a sorted list of links related to arxiv with titles and dates.A webpage displays a link library interface with a search result for hacker, showing one link titled Searching for Suzy Thunder from theverge.com dated 2020-01-22.

Author Margaret Atwood plays with Claude and reports back on her experience.

Interesting read: NYT is using a custom LLM tool to track trends within the “manosphere,” as reported by the Nieman Journalism Lab.

A friend of the blog told me a story about a Substacker who uses AI to summarize books and then publishes AI-generated content about those summaries, never reading the books herself, and yet has a ton of followers. I’d guess at least some of those are purchased, betting that a high follower count will beget more followers by suggesting clout and credibility she didn’t earn as a reader talking to fellow readers. And followers aren’t subscribers, but that’s the business bet.

People are lookie-loos, they get curious when something is doing numbers and creating activity, so inflating follower counts is a real and persistent strategy. None of this is new. But best practices still hold regardless of which technologies you layer on top. Marketing erodes trust when it prioritizes short-term gains over honesty and reliability.

It’s strange to live in a time when you can’t reliably distinguish someone who has engaged with ideas from someone who automated the appearance of engaging with them.

Anecdotally hearing about LLMs being weaponized in divorce and custody, including inundating the other party with slop to drive up the opponent’s legal fees. Worse, the sycophancy is tuned to and confirms the aggrieved party’s grievances, regardless of their real-world relevance in court.

Through a new quiz, NYT asks readers to rate passages of writing against AI. Despite thinking I could spot the AI writing, my results were 50/50.

AI in practice: Chatbot tool comparison

Come look over my shoulder while I explore how and whether LLMs are good writing tools: Here’s a wee version of the LLM comparison exercise I did with my team. We’ll make it a two-fer so you can see how the “good writing” skill works in practice, though we’ll see how that actually goes.

One of the more useful things you can do with an LLM is hold up a few ideas side by side and apply lenses to them. I know this history pretty well, so I asked a series of LLMs, why is Wisconsin’s cultural identity and cohesion stronger than Indiana’s, from a historical and business perspective?

Here are the answers in one doc, for comparison.

Each LLM will give us more or less the same story, different flavor. Within the industry, the differences across the models reflect “model personality.” Asking “why” instead of “whether” will probably drive the answer to favor Wisconsin. Using multiple lenses (two states, historical + business, identity + cohesion) forces the LLM to cross-reference across more of its training data, which tends to produce a more comprehensive answer.

For all the chatter about consciousness and whatever, remember that an LLM is an infinite series of if/then/elses applied to human language and semantics, so being able to talk about language and communication, getting meta with the tool and how you think through language, helps a lot when using one. This is maybe the one thing I like about experimenting so hard with the tools. I’m thinking about the technical side of writing and enjoying it quite a lot.

Functionally: all of them acknowledge hard historical truths within the subject matter and don’t shy away from critical perspectives, which is good. Both Gemini and Copilot include in-line links, which lets you judge the output’s authority in the moment as a reader. I liked Copilot’s more than I expected here. Claude’s answers are more lyrical and do provide more context, and yet do not encourage checking against outside sources by providing links within the output. And you can see that even with the good writing skill calling out hard bans on certain structure, Claude plows right through them.

Model personality: Claude favors sociological answers to Copilot’s economic answers. Claude is also highly intellectual and narrative by comparison, and that narrative style can mask nuance by sinking relative context within the storytelling. Gemini simplifies, boosts and cheerleads where the others don’t, and really goes hard on Wisconsin’s reputation as a drinking and Packers state when there are stronger structural arguments in play. Copilot is tricky because it looks authoritative like a briefing, which also makes it easily “extractible” for the user, but every citation requires authentication unless this is one of those “good enough” tasks.

As a writer, something I find annoying across the whole spread is the semantic reveal. LLMs are semantic machines, and it is persistently revealed in ways that are weird to the human ear. All of them go out of their way to describe things as “structural,” “connective” as in “connective tissue,” “load-bearing” and “legible.”

Finally, I included a second tab where I asked Claude for analysis across the four outputs, where it suggests that my framing of the question is altogether kind of problematic. It shows how a strong prompt is sometimes also a bad approach.

There are a lot of possible takeaways here, but I’d rather set aside the question of which tool is “good” or “bad” or “better” and think more about the patterns across the tools and their implications.

One of the tricky things about consumer AI tools like Claude and Gemini is that the experience varies widely depending on the person using it, and it’s not always clear why. I have spent a lot of time learning the tools so I can advise on them in my work, and this variance of experience has become a frustrating part of the deal.

I manage a team of writers and creatives at work, and we are expected to be familiar with the tools, despite complex and sometimes hostile feelings about the political and environmental implications of this sector. That’s quite a pickle, organizationally, managerially. Borrowing from Haraway, I thought, okay, what if we take these tools seriously as a team of writers and creatives and put our professional standards up against them?

Among other exercises, I did a couple of comparisons on my team that help create discussion around the “plausibility” question. People dismiss LLM outputs as being merely plausible answers, rather than accurate or factual ones. And that’s correct; they are, and that’s the design. In many cases, plausibility is fine. Take Wikipedia, for example, which we understand to be a pretty good source, a plausible source, unless you’re writing a formal paper requiring original sources.

I digress. Ultimately, we needed to understand together that LLMs are not a WYSIWYG tool and talk through the implications.

I asked everyone to run the same paper through their LLM of choice, prompting it for a plain language summary. We then copied and pasted it into a shared doc, and compared and contrasted for discussion. Upon discussion, we had several takeaways, including that they were all similar in spirit but sometimes varying wildly in style and approach.

Knowing that algorithms are responsive and not static, we did it again later in the day and copied and pasted our outputs into the shared doc. We compared and contrasted the difference between AM and PM. Again, it was similar in spirit but varied in style and approach. Some changed dramatically. One team member whose morning summary had been jokey and conversational received a much more staid and serious version in the afternoon.

At the time, I asked Claude to explain the variance: “Even with the same prompt and source material, LLMs don’t produce identical outputs each time. This is by design — there’s a degree of randomness (called “temperature”) in how the model selects words, which means each run produces a slightly different path through the text.”

Anyway, this got our gears turning on how (and whether) to approach LLMs as a team and as individuals and led to good group discussion. (It’s important to create space for criticism and critical approaches here.) It also gave us more confidence as a team responding to this new layer of complexity in our work, and helping our professional contacts and peers think about how to approach the tools and when and whether to use them. There will be tasks where AI-based tools are “good enough,” and tasks where they are not.

The swirl of mystery and speculation around this sector has people up in arms, and it’s useful to have approaches that give people firsthand experience and to see how the experience works for others. The god trick of the singular interface turns out to be a bear for navigating it in the workplace, where our work is foundational, prosocial and specific.

Good and bad uses of AI on a website, from Sumy Designs.

Applying a Claude writing skill

LLMs have a default house writing style with identifiable patterns: sentence fragments for emphasis, “not X, but Y” constructions, lots of hard contrast, atmospheric openings, heavy use of em dashes, and heavy use of marketing language. This reflects the semantic construction of an LLM. Custom instructions can override these defaults. A custom skill is a set of instructions within your account that modify how the model generates text. When you paste instructions into your profile settings, Claude reads them at the start of every conversation and adjusts its output accordingly.

I began using Claude daily for light writing tasks about six months ago, and over that time I started cataloging the patterns I was consistently editing out, including the terrible “not X, but Y” construction that showed up in nearly every response, and persistent em dashes used as all-purpose connectors when other punctuation is more appropriate.

I went through several iterations of bullying Claude into submission, narrowing the scope each time, before arriving at this version, which focuses specifically on writing mechanics and hard prohibitions.

You’ll need a paid Claude plan (Pro, Max, Team, or Enterprise). Free-tier accounts don’t have access to custom skills.

• Within the app, navigate to Customize > Skills and Create new skills
• Select add a new skill and Write skill instructions
Copy and paste the copy from this file into the skill, making note of the name and description boxes. Feel free to tinker.
• Save your changes.

Note: The instructions in the linked file are Claude’s work, not mine. They came out of months of conversation, where Claude would analyze my style notes, and the file evolved from there. They read a little strangely because of that process. If I’d written them from scratch, they’d sound different. But looking at the file you can see what Claude responds to and how it works.

Claude will apply these instructions to every new conversation going forward. Existing conversations won’t pick up the change, so start a fresh chat to test it. If and when Claude struggles to apply the skill, call it out specifically in the prompt, such as, “Revise this for length using the good writing skill.”

The skill specifies constraints in a few categories and the instructions are plain text. As you go, you can also ask Claude to analyze previous conversations for suggested additions to the skill, which Claude will produce and implement within the chat. Each rule operates independently, so removing one doesn’t affect the others.

Claude processes custom instructions at the start of every conversation, before it generates any output. The instructions function as constraints on the model’s default behavior. The model doesn’t always follow every instruction perfectly and the results vary by task. You will still need to edit.

WaPo on the many issues of LLM house writing styles. I have much to add. Here are some notes on applying a writing skill to override house style on Claude.

This new report from Anthropic is depressing at best, as it tries to measure which employment sectors carry the most exposure around AI expansion into the economy. In short, the tech is likely to impact two groups the hardest: educated professional women, and young workers for whom the career ladder will never materialize. In a right-side-up world, this would change the political dynamics of any policy response considerably. In this one, I don’t know.

Anthropic’s positioning here is curious, very god tricky. They are claiming the mantle of responsibility and transparency while predicting an inevitable end nobody wants, that they’re also selling as a service.

I still think much of the forecasting is oversold – the tech performs well in optimized environments, and last mile issues are a perennial concern in any engineering venture because the practical world is non-optimal. Time will tell, and there are big incentives in play. But the hunger and animus around the forecast feel bad.

Apparently one thing LLMs excel at is deanonymization at scale. The original promise of pseudonymity online was social and normative, over and above any question of technical depth: decent people don’t try to unmask you, because why. What strikes me today is how what used to be unacceptably antisocial behavior online is now both automated and unremarkable.

Over the last couple of weeks, I asked a couple of chatbots what could be known about me from this pseudonymous site, where I am more intentional about what I choose to reveal and conceal. It pulled the obvious but also drew conclusions based on a few geographic points I’d made in context that were both revealing and correct. I also noticed that it only drew from the top two pages of information - anything beyond page two of posts wasn’t part of the compute. Archives are for humans?

People assume that there is some computer magic on the backend where the LLMs connect all your account logins behind the scenes, but no, in fact it does all this through inference, by linking your digital trail, your friends, your breadcrumbs of likes and hearts and follows, and obvs your posts, into a picture of who you are, practically and demographically.

Second wavers in tech and engineering talked a lot about the “god trick” of presenting knowledge and information, particularly around math and science, in a way that suggests its objectivity is eternal, immortal, unknowable. Work by Safiya Umoja Noble and others extended this lens to Internet search and architecture, finding that the search algorithm was never neutral, instead it was a series of business decisions wearing neutrality like a costume, creating a customer service experience. LLMs take that same trick and compress it further.

It’s an old idea, and one I’ve been drawing from while I tinker with Claude, which is purportedly the best in the game. The “god trick” is baked right into the AI interface: one input, one output, an authoritative-seeming answer, offered without named perspectives behind it, trained on text produced overwhelmingly by a narrow demographic who has historically had access to both literacy and publishing, by programmers and new media drawing from the same well. Smushed together, it gives the impression that consensus exists where there are in fact many, many loose ends.

I increasingly find it annoying that even “good” AI outputs seem fixed on phrases like “key,” “core,” “exist,” “actually,” “never,” and possibly the worst sentence structure of all time, “it’s not X, it’s Y” — and I’ve begun to recognize how LLMs work like autocorrect for phrases and ideas, drawing from ranked search sources first before fanning out to more obscure sources, trying to determine and assert what’s important to me, a user known by demographics and data. It feels like a big linguistics machine, which is pretty cool in some regards, but also aggressively semantic. The math doesn’t always work to connect me to what I want to find because I am situated in my individual context in ways LLMs are not able to understand, with my memory, in my body, with my unique experiences, which shape and translate meaning for me as I interact with the world (and the web).

And so for you, in your body and memory and experience. An LLM can approximate the outputs of an experience without having access to the experience itself. Sometimes this is useful, sometimes it’s reckless.

Overall the dynamic reminds me of the famous scene from Good Will Hunting: Claude is a smart kid, and he’s never been outta Boston.

I’ve been running a series of informal experiments with my team to understand how we should advise people on using AI tools like ChatGPT, Claude and Gemini around our work.

In our last session, I gave everyone the same task: ask your AI tool to summarize a single article in plain language. Every person got a factually accurate summary, but the style and depth of analysis varied widely from person to person. We tested this further by entering the same prompt multiple times. Each time, the output was slightly different. We copied and pasted everything into a shared document so we could compare side by side.

Later that day, I asked everyone to reopen the tool they’d used that morning and repeat the prompt that had worked best: summarize this article in plain language. The outputs shifted again. Some changed dramatically. One team member whose morning summary had been light and conversational received a much more formal, serious version in the afternoon.

I then asked Claude to compare the morning and afternoon versions of one summary. “In short, the morning version emphasizes practical implications and accessibility; the afternoon version prioritizes completeness and specificity. They complement each other well — someone reading both would get a fuller picture than either one alone provides.”

Why would Claude produce these differences despite the same source material and prompt?

I asked the source.

Even with the same prompt and source material, LLMs don't produce identical outputs each time. This is by design — there's a degree of randomness (called "temperature") in how the model selects words, which means each run produces a slightly different path through the text.

But the differences here go beyond random word choices. The morning version landed on a journalistic structure while the afternoon version landed on an academic one. That likely reflects what the paper calls sensitivity to framing and order effects — one of the cognitive biases it catalogs. Small, invisible differences in how the model’s internal processing unfolds at the start of generation can cascade into meaningfully different organizational choices. Once the morning version opened with a news-style framing, the rest followed that pattern. Once the afternoon version opened with a descriptive academic tone, it followed that thread instead.

This is actually a concrete example of the robustness issues the paper identifies. The underlying reasoning about the paper’s content is consistent across both summaries — neither gets the facts wrong. But the presentation varies in ways that weren’t requested, which means someone relying on just one version would get a subtly incomplete picture depending on which run they happened to see.

It’s a useful reminder that LLM outputs are one plausible version of a response, not the definitive one.

It’s important to understand how these tools work. What looks like an authoritative answer is really just one of many possible outputs the system could have generated. The text will sound plausible whether it’s accurate or not, and you may get different results each time you ask.

I asked Claude.ai for a plain language summary of this buzzy new research paper. Claude offers the following:

Researchers from Caltech and Stanford conducted the first comprehensive survey of how and why large language models (LLMs) — AI systems like ChatGPT or Claude — fail at reasoning tasks. Rather than focusing on what AI can do well, they cataloged and organized the ways it goes wrong.

Why it matters
LLMs are increasingly used in high-stakes settings, but their reasoning failures are often unpredictable and poorly understood. The researchers argue that systematically studying failures — like accident investigations in aviation — is essential for building more reliable AI.

The researchers created a framework with two dimensions:
• Type of reasoning: informal (intuitive, social), formal (logic, math), and embodied (physical, spatial)
• Type of failure: fundamental flaws built into the architecture; domain-specific weaknesses; and robustness problems (inconsistent performance when small details change)

Informal/intuitive reasoning
LLMs exhibit human-like cognitive biases — confirmation bias, anchoring, framing effects — but without human ability to recognize and correct for them. They also struggle with “theory of mind” (understanding what others believe or intend), and with applying consistent moral or ethical reasoning.

Formal/logical reasoning
LLMs often can’t reverse simple logical relationships (if they know “A is B,” they may not infer “B is A”). They struggle to chain multiple reasoning steps together. Basic counting and arithmetic fail in ways that seem surprising given their other capabilities.

Embodied/physical reasoning
LLMs have poor intuitions about the physical world — gravity, spatial relationships, object properties — because they’ve learned only from text, not from physical experience. This extends to visual AI systems as well.

Many failures trace back to how LLMs are trained: they predict the next word in a sequence rather than reasoning deliberately. This makes them good at pattern-matching but unreliable when tasks require genuine logical inference, especially under slight variations in how a question is phrased.

Researchers have proposed fixes including better training data, techniques that force step-by-step reasoning (like “chain-of-thought” prompting), connecting LLMs to external tools like calculators or physics simulators, and architectural changes. However, no single fix is comprehensive — many improvements in one area don’t transfer to others.