I generated roughly 180,000 words of LLM-assisted writing last year. I deleted all of it. I have since started using LLMs again, but for different purposes, in different ways, with a clearer sense of what they can and cannot do.

This is what I learned from the experiment, what I deleted, what I kept, and how I use the tools now.

The first experiment

I started using LLMs for writing in early 2024. The first uses were utilitarian — drafting emails, summarizing documents, generating lists. The uses expanded as the models improved. By mid-2024, I was using LLMs for longer-form work: blog posts, essays, the first draft of a book I was working on.

The process I developed was, I thought, sophisticated. I would outline the piece myself. I would write a detailed prompt describing what I wanted — voice, structure, examples, length. I would generate a first draft. I would edit the draft heavily. The final product would be a hybrid.

The hybrid was faster than writing from scratch. I estimated I was producing roughly twice as much output per hour. The hybrid was also worse. I did not realize how much worse until I deleted it.

Why I deleted it

The reason I deleted it was not that the writing was bad in any obvious way. The drafts were competent. The grammar was correct. The structure was logical. The arguments were defensible. On any individual page, the writing was as good as my own.

The reason I deleted it was that none of it was mine.

This is hard to articulate without sounding mystical. The drafts were not mine because they were not built from my specific experience. They were built from the training distribution of the model, which is to say they were built from the average of millions of other writers. The drafts were the average of other people's writing. Mine were missing.

The specific things that were missing:

The specific examples. When I write about moving apartments, I write about the 19th arrondissement and the radiator that never worked and the neighbor who played piano at midnight. The LLM draft had generic examples of moving. They were plausible. They were not mine.

The specific opinions. When I write about cooking, I write about the Mark Bittman cookbook that started it for me. The LLM draft had opinions that were common across food writing. They were correct. They were not the opinions I have arrived at after ten years of cooking.

The specific voice. The LLM draft had a voice, but it was the average voice of internet writing. It was the voice of the training data. My voice is not the average voice of internet writing. The hybrid was a slightly better than average voice, applied to my subject matter.

The specific mistakes. This is the most counterintuitive one. My own writing has characteristic mistakes — I overuse certain constructions, I have a tendency to parenthetical asides, my paragraph lengths vary. The LLM drafts had none of these mistakes. The drafts were smoother than my own writing. Smoother is not necessarily better. The smoothness was the signature of the average voice. My mistakes are the signature of my voice.

The accumulated effect was that the LLM drafts felt like an approximation of writing. The approximation was good. It was not the thing.

What I tried to fix it with

Before deleting everything, I tried to fix the issues. I tried more sophisticated prompts. I tried giving the LLM samples of my own writing. I tried post-editing more aggressively. I tried restricting the LLM to specific sections.

The result of all of this was a kind of writing that was technically mine — I had edited every sentence — but that did not feel like mine when I reread it. The voice was a hybrid. The hybrid was competent. The hybrid was not what I wanted to publish under my name.

The deeper issue was that the LLM was doing the parts of writing that I most valued. The structure of the essay. The choice of examples. The rhythm of the sentences. These are the things that make writing feel like writing. Having the LLM do them felt, in retrospect, like outsourcing the reason I write.

I do not write to produce text. I write to think. The thinking happens in the specific examples, the specific opinions, the specific voice. The LLM drafts were produced by a process that did not include my thinking. The drafts were technically mine. They were not generated by me.

What I kept

Not all of the LLM-assisted writing was deleted. The pieces I kept:

Outlines. I still use LLMs to generate outlines. The outline is structure. The structure is something I am happy to outsource. I take the outline, edit it to match what I actually want to say, and use it as the skeleton for the piece.

Research summaries. I use LLMs to summarize long documents, articles, and research papers. The summary is the input to my own thinking, not the output of my thinking. The line is clear.

First-pass editing. I use LLMs to check my own writing for grammar, repetition, and structure. The LLM is a better proofreader than I am. The content is mine. The proofreading is the LLM's. This is a clear division of labor.

Headline generation. I use LLMs to generate headline options for my own articles. I almost never use the LLM's headline. The process of generating options helps me think about what the headline should be. The output is mine.

Code. I use LLMs for code. Code is a different category from prose. Code has clear right and wrong answers. LLMs are good at code. The writing-about-code I do is my own. The code itself is collaborative.

Translation. I use LLMs for translation. The translation is not literary. It is functional — getting the meaning across. The literary translations I do are my own or I hire a human translator.

The pattern is that I use LLMs for tasks where the output is a tool, not a product. The output of outlining is a tool I use to write. The output of writing is the product. The product should be mine.

How I use them now

The current setup:

Drafting. I draft by hand. The drafting process involves writing the first paragraph three or four times until it sounds like me, then writing the rest with the first paragraph as the voice calibration. This takes longer than LLM drafting. It produces writing that is mine.

Editing. I edit by hand first. The first edit is structural — does this paragraph belong here, is this example the right one, is the argument coherent. The second edit is line-level — sentence rhythm, word choice, the specific construction. Only after both passes do I use an LLM for a final proofreading pass.

Fact-checking. I use LLMs to check facts, dates, and attributions. The LLM is faster than Google for these lookups. The LLM is also sometimes wrong. I cross-check anything important.

Translation. I use LLMs for first-pass translation. I edit the translation to make it sound natural in the target language. The translation is collaborative, with the LLM doing the bulk of the work and me doing the voice work.

The cumulative effect is that I write less than I would with LLM drafting, but what I write is mine. The output per hour is lower. The quality per word is higher. The output that gets published feels like it was written by a person, which it was.

The broader question this raised

The deeper question I have been thinking about since the experiment is what writing is for.

If writing is for producing text — generating words, paragraphs, articles, essays — then LLMs are a productivity win. The output is faster. The output is competent. The output is publishable. The productivity gain is real.

If writing is for thinking — figuring out what I actually think, working through an argument, building a position — then LLMs are a productivity loss. The thinking is what I outsource when I outsource the drafting. The text I produce may be publishable. The thinking I do is less than the thinking I would do if I wrote it myself.

The distinction matters because the writing I keep coming back to — my own and others' — is the writing that was the product of someone's thinking. The writing I remember is the writing where the author was working something out. The writing I forget is the writing where the author was generating text.

LLMs are very good at generating text. They are not good at thinking. The writing they produce is the writing that is forgotten.

When LLMs are the right tool

A handwritten notebook open on a desk beside a laptop, the contrast between manual writing and machine output

A notebook with handwritten pages next to a laptop on a wooden desk

LLMs are the right tool for:

  • Functional text (emails, summaries, listicles)
  • Repetitive content (product descriptions, SEO copy, social media posts at scale)
  • Code and technical writing
  • Translation
  • Research assistance
  • Editing and proofreading

LLMs are not the right tool for:

  • Personal essays
  • Memoir
  • Literary fiction
  • Argumentative essays where the argument is the point
  • Any writing where the voice matters

The categories are clear. The problem is that most of the writing people want to do is in the second list. The first list is the writing people do for work. The second is the writing people do for meaning.

The honest summary

I deleted 180,000 words of LLM-assisted writing because none of it was mine. I have since started using LLMs again, but for tasks where the output is a tool, not a product. The result is that I write less, publish less, and produce work that I am willing to put my name on.

The technology is extraordinary. The technology is also not me. The writing that gets published under my name is the writing I wrote. The technology helps with the parts I do not value as much. It does not replace the parts I value most.

If you are using LLMs for writing and the output feels wrong, it is not you being precious. It is the technology not being designed for what you are trying to do. The technology is designed for functional text. The writing you care about is probably not functional text.

The tools are good. The tools are not for everything. The art is in knowing the difference.