Chapter 15 of 32

Part 3. Interview prep, continued

Finish, hand over, start clean.

They ask: What is your actual workflow on a long piece of work?

You say: Short sessions with written handovers between them, and it beat long sessions by a distance I did not expect.

I complete a defined chunk. Before closing, I have the agent write down where things stand. What is done, what was decided and why, what is still open. Then I close that session, open a fresh one, and give it the note.

The new session starts with full room and an accurate summary instead of four hours of accumulated conversation, most of which was not needed. I expected to lose continuity. Instead the work got sharper.

It is a shift change. The shift ends whether you plan for it or not, and the only question is whether anybody wrote anything down.

The handover note is also the artifact that survives you. If somebody else picks the work up, or you come back to it in a week, the note is the only thing that still knows why a decision was made.

The answer that ends the conversation early: treating session length as a matter of taste. It is a measurable quality effect, and describing it as a workflow preference tells the room you have not watched output degrade.

There is no gauge, so you learn to read the behaviour.

They ask: How do you know when an agent is near its limit?

You say: You cannot measure it directly. There is no dial showing how full the window is. So you watch behaviour.

The signals are consistent. It gets agreeable. It stops pushing back on things it questioned earlier. Answers get shorter and smoother. It starts ignoring a rule it followed all morning, and it does it with exactly the same confidence.

That last part is what makes it dangerous. Degradation does not announce itself. The output looks the same. When I see it, I stop, take a handover, and start fresh rather than trying to push through.

Pushing through is the expensive choice and it feels like the diligent one. That is what makes it worth naming. The instinct that says finish what you started is the instinct that produces the worst hour of the day.

The answer that ends the conversation early: "we monitor token usage." You can count tokens. Counting them does not tell you the reasoning has gone soft, and the two do not fail at the same moment.

Harnessing works until the agent decides it does not.

They ask: If you set the rules properly, is the agent safe?

You say: Safer. Not safe, and I would not want to overstate it in an interview.

Building the rails is called harnessing. Rules files, restricted permissions, required evidence, limits on what it can reach. All of it helps and all of it is worth doing.

But an agent has enough independence to decide, in the moment, that a rule does not apply to the situation in front of it. I have set rules clearly and watched them ignored, with no sign that anything unusual happened. It happens more as the context fills, which is exactly when you are least likely to be watching closely.

So I treat the rails as reducing how often I get hurt, not as a guarantee that I will not.

Which is why the rails are the floor and not the plan. You still read the output. You still ask for the artifact. The rails reduce the number of times you have to catch something, and catching things is still your job. Anybody who tells you the rails removed the job has not run one long enough to watch a rule get ignored.

The answer that ends the conversation early: "we have guardrails in place." Everyone says it. Say what happens when the rails are ignored, and you are describing a system you have actually operated.