↓ Skip to main content

Vibecoding, prompt engineering, or just good old fashioned collaboration?

October 7, 2026 · Jared Knowles

There is a lot of discussion about prompt engineering that makes it seem like there are these magic incantations you need to learn and practice to get high-quality LLM results. The record from this project suggests a different lesson — the importance of review and engagement.

In my previous post on agentic data science, I focused on quantifying the work in tokens and hours. As I promised in that post, this post explores the qualitative side of the work — specifically, I’m interested in evaluating my own work with the LLM and answering: how often did the model make mistakes and how were those mistakes caught?

This evaluation of my own process led to four conclusions:

  1. A good plan to ground each session with the AI keeps the project moving, but moving beyond the plan in the same session brings risks.
  2. Stay skeptical when the AI is proposing designs for human interfaces to the data. I should have asked more questions before accepting proposed defaults.
  3. Ground the project in published documentation and data. The AI agent can keep its sessions grounded in this truth and avoid both of us speculating about why the data looks a certain way by going straight to the source for the answer.
  4. Recording sessions and periodically reviewing them is worth doing!

For this evaluation, I am using my project session record: 74 sessions in which I was at the keyboard, from April 16 to September 4, with 717 messages from me and 16,735 from the agent. (The 569 sessions that the last post counted include the agent’s own subagents; those have no person in them.) Using AgentsView (opens in new tab), I searched every agent message for an admission — you’re right, my mistake, I introduced — and every message of mine for a correction, a question, or an interruption, and then instructed the LLM to read and summarize each exchange.

This process identified 75 distinct errors the agent made. Four of them reached users, including the two API outages described below. The rest were caught before deploying. The table below shows how errors were identified and which were caught before production. Nearly two-thirds of the errors were caught by an agent: 27 by the agent that made them and 21 by a second agent reviewing its work. I caught 12, tests and build checks caught 11, and four reached users.1

How the error was caughtDistinct errors
The agent’s own verification, in the same session27
A second agent reviewing the first21
Me, by domain knowledge or a direct question12
A test, build check or CI (continuous integration) failure11
Reached users before anyone caught it4
Total75

Where the mistakes were

The agent and I made mistakes throughout the project, as the figure below shows. The mistakes are distributed pretty evenly across time. Like in the previous post, I attribute some of this to my process for working with the LLM improving as the work became more complex and detailed in the latter half of the project. Below I will get into the transcripts and share my hypotheses for where and why agent mistakes occurred.

Human turns per interactive session, in order. 74 bars from April 16 to
September 4 2026; half of the sessions take between five and 15 turns and
the tallest takes 38. Filled circles mark the 27 sessions where the agent
admitted at least one of the 75 errors, open squares sessions with an unclear prompt
of mine. Both kinds are spread across the whole project.

The importance of starting a session with a plan

An agent session will run unattended for a long time if I start the session with a one- to three-page handoff or brief, usually written by the previous session’s agent and approved by me. Measured in terms of how long before I need to weigh in, the agent does nearly four times as much work after a brief as after a typed command: a median of 36 messages, 38 tool calls, and 24 minutes.

However, sessions that open with a brief end up about the same size as sessions that open with a typed line. In both types of sessions two-thirds to three-quarters of the agent’s work comes after the opening run, steered by me. The table covers the 68 sessions with at least 10 agent messages; the other six have fewer than 10.

Opened with a briefOpened with a typed line
Sessions2840
Median agent messages per session254250
Share of the agent’s work inside the opening run32%26%
Median active hours per session2.62.2

Beyond the brief is where trouble begins

For each session, I measured how much of the agent’s work came before and after my first reply to the opening instruction.

The longer I kept steering after the opening run, the more often the agent admitted an error. Sessions in which I sent ten or more messages after the opening run hold half of the agent messages in my sessions and 73 of its 98 admissions. Sessions in which I sent three or fewer hold 15% of the work and two of the admissions. The figure below shows this pattern.2

Share of the agent’s work and of its admissions, for sessions grouped
by how many messages I sent after the opening instruction’s run: 0–3 messages
(15 sessions), 4–9 (27) and 10 or more (26). The share of work rises from 15% to
35% to 50%; the share of admissions rises faster, from 2% to 23% to
74%.

There are two things happening here. First, I was not disciplined about managing the context each session was working with, and I let sessions run too long. But the second is more interesting. I ignored signs that the session was at a good point to stop, slow down, and be more planful. After half of my briefs (20 of 40) the agent stopped to ask me something; after one in eight of my typed prompts it did. The agent tends to reach the end of the brief and start asking questions when it gets beyond the boundaries of that brief.

I have learned that, when this happens, it is time to a) start a new session, b) write a detailed plan for what work comes next, or c) ideally, both.

API outages in production happened when I pushed beyond the plan

Both times the API broke, the session started with a clear brief and the agent stopped after completing that work. For example, the second time, the session started with a brief for a local rebuild of one year of data that said “Do NOT publish.” The agent finished, reported, and stopped. Instead of reviewing the work, I simply asked “Where is 2024?” and told it to commit and keep going until the 2024 ingest succeeded. I did not write a brief, plan, or rules for this extension of the work. From those two lines, with nothing more from me than “status?”, one answer to a question, and “Resume” twice, the agent sent 408 messages over the next 17 hours, stopping overnight until I told it to resume. I hit publish on the 2024 data the next morning and the API stopped answering.

Panicked, I asked what happened and the agent said:

Agent

That’s my error, and it took the API down. I told you uscogdata gates on schema_version >= 5 — I read that from R/basis.R, which governs basis resolution. The actual boot-time check is a separate function with an explicit allow-list.

But I had failed to check or ask an agent to check this. In fact, I had let my eyes skip right over this message:

Agent

uscogdata gates on schema_version >= 5 (a floor, not an equality pin), so v7 satisfies it — no reader breakage “Report before Task 13 (publish) — owner-gated, do not run”.

I was the “gate” and I didn’t do my job!

Similarly, the first outage I created began with the same structure: three small fixes came up mid-session. The agent warned me, before I merged, that one of the fixes removed something other code might depend on.

Before I merged:

Agent

#16 removes a capability the issue didn’t mention.

After the outage:

Agent

I flagged the capability loss on uscogdata#16 and checked uscogdata’s callers — I did not grep this repo for them.

Project organization could be improved

This sequence highlights another problem: both outages occurred when work crossed from one repository into another. The pipeline, the R package that reads its output, and the API that serves it live in separate repositories, and each outage was a change in one that broke the next one downstream. The API ran no tests before a merge in the first outage and built against whatever version of the reader package was current, so a change to the reader reached users without any change to the API.

My mistake was that I extended the work in a line or two without engaging in any planning and without being mindful of which repositories the work would touch. Meta-planning, which agents are quite capable of leading you through now, helps avoid these kinds of mistakes. In my projects since, I’ve started using a custom-built plugin called “compass.” I start and end every session with a slash command that keeps a decision record and a journal of the work, and stores all those plans in a place any fresh agent session can quickly find. This slows things down, but it also leads to fewer “YOLO” moments and fewer breaks reaching production. I also think more carefully now about how I organize projects that are modular with more than one repository.

When I give bad instructions, the session spins its wheels

When a typed prompt of mine is ambiguous, the agent’s first move can be to go and look rather than guess. The cost is small, but it can add up over time. On July 25 I wrote:

Me

Can you start on #2 while the build is working?

The agent first checked whether a build was running (it was; it had not started one), then fetched two issue descriptions, then asked:

Agent

“#2” is ambiguous — which did you mean? – uscogdata #3 — gov search bug (my step 2): My ranked list had ‘uscogdata #3 — gov search doesn’t work’ as step 2, and ’the build’ as step 1 — so ‘start on #2 while the build is working’ parses as ‘do step 2’. […] – uscogdata #2 — households denominator: Read literally as an issue number: ‘Extend population data to be households as an alternate spending denominator.’ A feature that adds a second per-capita basis alongside population.

Seven minutes and 34 tool calls just to ask me to clarify whether I meant the bug. I should have just said so to start!

The same thing happened on a larger scale on July 10, when I wrote that to finish the dataset “we need to finish the pipeline work.” The agent read the state of every stage of the pipeline before answering:

Agent

The pipeline is much further along than “incomplete” implies […] So the corpus isn’t “broken” — it’s a deliberate v0.1 slice. Which means “finish it” needs a definition.

It offered me four definitions of “finish it.” I picked three, and those three became the next month of work. Because I didn’t specify, the agent had to review the whole project to ask me to clarify. Time and tokens thoroughly wasted.

Asking the agent to justify a design pays off

When I wanted the work to continue, I usually said so in a word or two: “Proceed”, “Resume”, “Continue”. Of my 167 messages of 25 characters or fewer, 10 were followed by an admission, about one in 17. Of my 309 messages longer than a line (over 120 characters), 39 were, about one in eight.2 Those longer messages often reflected my need for clarification and more often led to sessions that uncovered an inconsistency or ambiguity that needed to be fixed.

The clearest case is a question about a function’s design. On August 3 I was reviewing the design for cog_balances(), the R package’s function for fund balances and trust holdings, before any of it was written. The proposed function took two filters, a subtype and a category. I asked:

Me

We need to talk through the intersection of the subtype and category parameters for cog_balances. Will a user ever use subtype = “general” and category = “Insurance Trust Balances” for example and what would each of these combinations mean?

The agent checked the data before it answered:

Agent

That question exposes a real design error.

For balances, each subtype sits inside exactly one category, so the two filters could only repeat or contradict each other. Of the 15 combinations, three were redundant and 12 returned nothing:

Agent

So to answer literally: subtype = "general" + category = "Insurance Trust Balances" returns zero rows for every government in every year. It isn’t a meaningful query, it’s a contradiction — and it fails by returning an empty tibble, which reads as “this government holds none” rather than “you asked an impossible question.”

The agent recommended dropping subtype, and the function shipped with category alone. An analyst who combined the two filters would have read an empty table as a government holding no money in those funds, with nothing to tell them the question itself was impossible.

There are probably places in this project where I should have asked such pointed questions and didn’t. It is important to stay engaged in the design and development, especially where the work will interface with users — what appears sensible to the agent may not work for users at all.

Where my judgment and expertise were needed

Four of the 40 briefs I sent ended with me interrupting. Two of those were me adding a thought within a minute of sending the brief. The other two came 17 and 73 minutes into work I had not wanted done the way the agent was doing it. On July 30 I pasted a 14,000-character issue into a session and the agent began implementing it. Seventeen minutes in:

Me

Don’t I need to help make some decisions about the recatgorization[sic] and how to handle these revenue concepts?

I did. The agent had already made one of them, replacing a test in a way that contradicted a ruling I had given earlier in the session, and the session stopped so the rest could be mine. Writing better specifications up front can head this off. Relying on the 14,000-character output of one agent to drive the next without giving it a review is a recipe for moving right past consequential choices.

Is this vibecoding?

The agent wrote the R code for this project — not me. I contributed the goal of unifying these datasets over time, validation procedures, and vision for how users will interact with this combined data. These came from my many years of experience and desire to make this data more accessible. I had done many mini-experiments coding a longitudinal series of a few specific counties, following a few specific categories of spending, and making some data visualizations for the most common categories.

These prior experiments and code allowed me to communicate my understanding of what is and is not in the dataset and the pain points in connecting it longitudinally. This helped me write good initial design requirements for the LLM and follow the progress of development. In this way I functioned a lot like a senior principal investigator on a research project with access to a team of talented data scientists with less subject matter expertise.

But that analogy does break down a bit because the agent was able to do something that was beyond my ability — read, index, and constantly and consistently reference the Census documentation to understand all the changes in definitions and collection rules. For this project there were about 1,350 pages of documentation. The agent was also able to cross-check our results against 15 public files with aggregate Census figures to help identify mistakes.

Could I have built this project without an LLM? Yes but it would have taken many more months. And the level of discipline it would have required to keep the details in focus is beyond what I could do alongside other professional responsibilities. In the presence of good documentation, the agent kept our work grounded in the realities of the data and the specifics of each collection.

The result was a collaboration. The agent made mistakes. So did I. I suggested design elements to push the project further (an API endpoint for LLMs specifically) and the LLM suggested ways to make the data easier to work with (giving users local bulk data replication via the R package).

Lessons I learned for current projects

I reviewed this project to better understand my own practice of working with LLMs. I learned a lot! I came away with a few things I’m keeping in mind in my work now:

  • I need to structure my work with more project management materials to keep the agent focused. Many mistakes the agent made came from reading the wrong data file, reading the wrong narrow slice of the code, jumping to conclusions, or testing things incompletely. These errors can be avoided with good rules and project structure so each agent session is operating with the same understanding of the project.
  • Gathering and incorporating data documentation into the project structure is critical. Any documentation and published data that exist can be used to check assumptions and verify calculations. This greatly reduces errors!
  • Writing detailed plans up front can unlock a lot of unsupervised agent time. I now often take two to three hours to develop a specification for a project or feature, carefully going back and forth editing a markdown file with the agent until it is very detailed and as unambiguous as possible.
  • I have a structured review process to deliberately pause and give the project a more thorough review. It is so easy to grow and expand the project with just a short prompt like “Where is 2024?” The agent will try hard to fetch and incorporate the data, and current models will make a whole host of decisions for you to avoid bothering you! Instead of doing this, when I find myself at this point, I work through a review with the agent of what has been built and evaluate the scope of what I would like to add.
  • Token management is key. I am often up against the weekly limits of my Claude Max subscription, so I have become very thoughtful about how to structure how the agent works: should we use subagents? Could a more affordable, less capable model do this work? What level of model “effort” is really needed to complete this task? I also try to plan up front to avoid the kind of mistakes described above that lead to wasted time and tokens.

What do you think?

I think of myself as a good R programmer with a lot of familiarity with the data and content in this project. But the result I was able to get and the speed at which we reached the goal still amaze me. In looking back on the project, I see there are ways I could have gotten this done even faster, for fewer tokens, and with fewer mistakes.

Partly, that is because the models are getting better. But just as importantly, there are tools to help us reflect. Working with an LLM is not like working with a Python interpreter — LLMs fail silently and even mask their mistakes (we do this as humans too). This makes it harder to learn how to use them better and is exactly why researching this post has been so helpful. It’s also why we heavy AI users can be so eager to talk about it.

So let me know what you have learned using LLMs. What has worked? What hasn’t? And what have you changed as you used them more? I’d love to learn from your experience as well.


  1. This kind of meta-analysis is really useful to understand how the tool is working and how to get the output you want. For this table, the count was built in three steps. First, a pattern search of every agent message in the project for an admission (you’re right, my mistake, I was wrong, I introduced, good catch and variants) returned 110 messages. Second, LLMs read the exchange around each admission against a written set of rules, set aside the 23 messages that admitted no agent error, and collapsed the rest into 85 distinct errors, counting an error once regardless of how many times it was restated, and recorded what caught each one. Third, I excluded 10 mechanical slips, such as a stray character or a broken code span, which leaves 75 remaining. The search only finds errors the agent admitted in those words, so an error it fixed quietly, reported in other words, or that nobody caught, is not in the table. The totals counted twice in one data release are one example. The count is a floor. ↩︎

  2. An admission is one agent message saying you’re right, my mistake or similar; the same error is often admitted more than once. An admission usually answers a message from me, so a session with more of my messages has more chances to show one. The chart shows where the errors came to light, not where they were made. ↩︎ ↩︎

Subscribe to The Civic Pulse

Get future posts delivered to your inbox.

Get The Civic Pulse delivered to your inbox.

← Back to Newsletter