AigentLab® No. 27 Download Guide

AigentLab® Field Guide  /  No. 27

Train
The Night
Shift.

Make. Grade. Keep the rule. No. 27
10.06.2026

Executive Summary

The Loop That
Grades Itself

Adapted from "I Built a Team of Self-Improving AI Agents". The creator runs agents on a nightly schedule. Each agent makes one piece of content, scores the piece against a written list of criteria, writes down what went wrong, and starts the next night from those notes. He wakes up to a better web page, a better short, and a better draft than the night before.

01
What a self-improving agent is

An agent with four jobs in one cycle. Make something. Grade the result. Write down the lesson. Try again next time with the lesson loaded. The grading and the note-taking are the parts people skip.

02
A schedule, not an always-on bot

The agents run while he sleeps. 1:00 a.m. web pages. 3:00 a.m. short-form videos. 4:00 a.m. Substack. 5:00 a.m. long-form video. Each morning he opens improved systems instead of blank ones. His core line: let agents work while you sleep, not while you work.

03
The origin: Auto Research

In March, Andrej Karpathy published Auto Research, an open-source repo built to let a model improve its own training code overnight. The post drew about 11 million views. The loop repeats four steps all night: change one thing, run a five-minute test, compare the score, keep the change if the score improves and discard the change if the score drops. One night tested about 700 ideas, kept about 20, and finished 11 percent faster than changes made by hand.

04
The concept transfers without the repo

You do not need Auto Research or GitHub. You need work you care about and a way to score the work. The same learning loop applies to a web page, a video, or a piece of writing.

05
Web pages first

Night one produced a clunky, linear page with a scrolling sidebar. Night two scored lower, because the agent had little training to work from. By night three the design, animation, and brand fit beat the best pages he built by hand. The rules the agent writes also govern pages he builds himself.

06
Why by-hand works, and what by-hand costs

Prompt, review, complain about colors or text, repeat. The back-and-forth produces good results because you bring the taste. The cost is your time in the chair. The loop hands the critique to an agent trained to judge the way you judge.

07
The criteria are the hard part

He calls the scoring criteria the most important and slowest piece to build. He found them by building pages by hand and noticing what he kept criticizing.

AreaWhat he checks
Contrast and fontsBold, legible type and enough contrast against the background.
SizingLarge elements, with room left for his face in the corner of the video frame.
ClarityNo vague "it" or "she" without a named subject. The page reads clearly to a viewer.
Visual craft and motionGlows, frosted panels, animated diagrams. More interest, not always more clarity.
Topic fitOne night produced a strong page about pi. Wrong channel, unusable.
Layout, type, interactionPlacement of text and graphics, big numbers, no generic serif fonts, clickable elements.
Restraint, wow, narrativeStrip what hurts the message, add one striking moment, keep a beginning, middle, and end.
08
One focus area per night

Each night the agent cycles to one of seven areas and improves only that area. Improving everything at once burns tokens and moves nothing far. When a run goes wrong, the agent writes itself a rule so the mistake does not repeat.

09
Beyond web pages

Short-form: a flat stacked graphic became a staircase animation with bounce, glow, and large background numbers, aiming for 10 percent better each night across hooks, scripts, and screenshots. Long-form AI-generated video is next. Substack runs on an agent trained on his Obsidian vault plus universal truths about titles, cover images, and hooks. Every system needs training by hand first.

10
The feedback dock

A markdown file where he leaves dictated notes each morning. The agent reads the feedback first and applies the notes before starting its own loop. He stays in charge without staying in the loop.

MAKE GRADE KEEPOR TOSS WRITETHE RULE NEXT RUN READS THE RULES FIRST ONE FOCUS AREA PER RUN
Fig. 01   The nightly loop, reduced to four moves

The takeaway

Set your taste once, in writing. The agent improves against your written taste every run, and your morning notes steer the next run.

New Video Concept  /  Core Message

Train The
Night Shift.

You fix the same mistakes in every AI draft because nothing you correct carries forward. Write your taste down as a scored rubric. Let an agent make one draft, grade the draft, and log one rule per run, on one focus area. Leave your notes in a file instead of a chat. Every run starts from the last run's rules, so your work improves overnight instead of during your workday.

Opening Premise

You use AI for the work a new professional ships every week: proposals, cover letters, outreach emails, case studies, one-page summaries. Each draft comes back about 70 percent right. You fix the same things every time. A vague opener. A pronoun with no subject. A paragraph too long. A price buried at the bottom. Next week you fix the same things again, because your corrections lived in a chat thread that ended.

The standard advice says write a better prompt. A better prompt helps one draft. The creator of "I Built a Team of Self-Improving AI Agents" hit the same wall with web pages: the back-and-forth worked, and the back-and-forth ate his day. His fix was a loop with a memory, run on a schedule while he slept.

Opening question

Look at your last three AI drafts. How many of your corrections were the same correction?

Teaching Points

01

Keep the loop. Skip the lab.

Karpathy's loop needed code, a GPU, and a five-minute benchmark. Yours needs four moves: make, grade, keep or toss, write the rule.

Pick one asset you produce at least once a week and judge mostly on quality, not facts: a proposal, a cold email, a portfolio case study, a LinkedIn post. Give the agent a fixed input it reuses every run, such as one past client brief. A fixed input means score changes come from the draft, not from a new brief.

MoveWhat you set
PickOne asset type, made weekly.
FreezeOne input reused every run.
ScoreThe same rubric every run.
CompareKeep the new draft only when the total score rises.

Rule

One asset. One input. One loop. Add a second loop only after the first runs clean for two weeks.
02

Write your taste before you hand it off.

The creator called the criteria the hardest and most important part. He found his by noticing what he kept criticizing. Do the same with your own edits.

Take your last five drafts. Next to each edit, write why you made the edit in four words or fewer. Group the reasons into five to seven areas. Turn each area into a question scored 1 to 5, with one sentence describing a 5. Write questions a stranger could score. "Sounds good" fails. "Names the client's stated problem in the first two sentences" passes.

Sample rubric: freelance proposal

AreaQuestion, scored 1 to 5
ClarityDoes every pronoun point to a named person or thing?
FitDo the first two sentences name the client's stated problem?
ProofDoes the draft include one real number from past work?
LengthIs the draft under 400 words with no paragraph over four lines?
RestraintOne offer, one price, one next step?
VoiceDoes the draft sound like you talking across a table?
Next stepDoes the close name a date and a single action?

Rule

No rubric, no loop. An agent grading against a vague rubric learns to please itself.
03

One focus area per run.

The creator rotates through seven areas, one per night. Improving everything at once spends tokens and changes little.

Score the whole rubric on every run, so you see whether a fix in one area broke another. Change only the focus area. Keep the new draft when the focus score and the total both rise. Toss the draft when either drops, and log why. Rotate in a fixed order: Monday clarity, Tuesday fit, Wednesday proof, Thursday length, Friday restraint, Saturday voice, Sunday next step.

Rule

Seven areas, seven nights. One area per run, the full rubric scored every time.
04

Turn every miss into a rule.

When a run scores lower, the agent writes one line into a rules file. The next run reads the rules file before writing a word.

Write each rule as an instruction with a reason: "Never open with 'I hope this finds you well.' Clients skip the first line." Specific rules work. "Be more engaging" does nothing. Cap the file at 25 lines. Once a month, merge duplicates, delete rules the drafts no longer break, and resolve any two rules pulling in opposite directions.

Rule

Paste the rules file into your own AI chats too. The night shift trains your day shift.
05

Leave notes, not chats.

The feedback dock is a plain markdown file. Each morning you read the overnight draft, dictate what bothers you into the file, and save.

Date every note and write the fix, not the feeling. "Too long" helps less than "cut the middle section to three bullets." Schedule one run per night at a fixed time, such as 1:00 a.m. ET. A laptop has to be awake and online for a local schedule, so use a cloud scheduled task if your machine sleeps. Expect a dip. The creator's second night scored worse than his first. Judge the loop after seven runs, not one.

Rule

Ten minutes of notes in the morning replaces forty minutes of edits in the afternoon.

Translation  /  The Night Shift, Built

Dana's Proposal Loop

Dana is a first-year freelance copywriter. She sends about six proposals a week to small business owners and spends about 40 minutes editing each AI draft. Every edit repeats last week's edits.

PieceSetup
AssetProposals for small business website copy.
Fixed inputOne anonymized past brief: a bakery owner wants a new About page and three service pages.
rubric.mdSeven questions from her own edit log, 1 to 5 each, 35 points max.
rules.mdEmpty on night one. 25-line cap.
feedback.mdDated morning notes, dictated in under 10 minutes.
outputs/One dated draft per run plus a score line. Nothing gets sent.
ScheduleNightly at 1:00 a.m. ET. Focus rotates Monday through Sunday.
BoundaryPractice brief only. Live client proposals stay manual until week three.
21NIGHT 1 19NIGHT 2 24NIGHT 3 25NIGHT 4 27NIGHT 5 27NIGHT 6 29NIGHT 7
Fig. 02   Planning estimate: rubric totals out of 35 over seven nights. Night two dips.

Planning estimate: editing time per live proposal drops from about 40 minutes to about 15 by week three, plus 10 minutes of morning notes. Log your own numbers for four weeks and replace these.

The nightly run prompt

Paste this into a scheduled task. Replace the bracketed parts. Keep the order, because the order sets priority: your notes, then the rules, then the work.

Nightly run You run one improvement cycle on my [asset type]. Work only inside the folder [path].
  1. Read feedback.md. Apply every open note first. Mark each note applied.
  2. Read rules.md. Follow every rule.
  3. Read rubric.md. Today's focus area: [area for this weekday].
  4. Read the latest kept draft in outputs/ and the fixed input in inputs/.
  5. Write one new draft. Change only what improves the focus area.
  6. Score the new draft and the latest kept draft on all seven questions, 1 to 5, with one sentence of evidence per score.
  7. Keep the new draft only if the focus score and the total both rise. Otherwise mark the draft tossed.
  8. If the new draft scored lower anywhere, add one rule to rules.md: one line, one reason. Never exceed 25 lines.
  9. Save the draft as outputs/[date]-[kept or tossed].md with the score table at the top.
Never send, post, or email anything. Never invent client facts, numbers, or quotes. If the input is missing, stop and write the reason to outputs/.
Morning feedback note[Date] Draft from last night. Fit: the opener talks about me before the client. Start with the bakery's problem. Proof: use the 30 percent bounce-rate drop from the dental client, not a vague claim.

Check yourself

Once a week, score one draft yourself before reading the agent's score. When your total and the agent's total differ by more than 4 points, rewrite the rubric question you disagreed on most.

Counterpoint  /  What the video leaves out

Before You
Trust the Score

A
A grader grading itself drifts

The agent scores its own drafts and learns what earns points from itself. Scores rise faster than quality. Spot-check against your own scores every week.

B
Three nights prove little

The win arrived on night three of a fresh setup, with no score scale or sample size on screen. Judge your loop over weeks.

C
Taste has no stopwatch

Karpathy's loop measured speed with a five-minute test. Taste needs a rubric to stand in for the stopwatch. A vague rubric produces a vague loop.

D
The bill stays off screen

Four nightly jobs burn tokens and the video never shows a cost. Cap runs at one per night per asset and log spend every Friday.

E
Rules pile up and collide

Every miss adds a line. Without pruning, the agent obeys contradictions. The 25-line cap and monthly cleanup carry the weight.

F
Off-topic output wastes a night

The pi page looked strong and served nobody. Fix the input and the topic list before blaming the design.

G
Your voice, still machine-made

Training on your vault narrows the gap. The draft stays AI-written. Disclose AI use where clients or platforms require disclosure, and send nothing unreviewed.

Closing Challenge  /  Seven Days, One Loop

Write The Taste.
Then Go To Sleep.

A new professional who used to fix the same five problems in every draft now reads a better draft each morning, leaves ten minutes of notes, and spends the afternoon on clients. Every correction made once stays made.

D1

Pick one asset. Pull your last five drafts and every edit you made.

D2

Tag each edit in four words or fewer. Group the tags into five to seven areas.

D3

Write rubric.md: one question per area, scored 1 to 5, one sentence describing a 5.

D4

Create rules.md and feedback.md. Run the nightly prompt once by hand and score the draft yourself first.

D5

Schedule the run at 1:00 a.m. ET with one focus area per weekday.

D6

Read the overnight draft. Dictate dated notes into feedback.md in ten minutes or less.

D7

Compare scores across runs. Delete one rule. Decide: keep the loop, fix the rubric, or kill the loop.

Make. Grade. Keep the rule.
Train the night shift.

AigentLab® No. 27

Take The Full
Field Guide

15 pages. The full rubric, the nightly run prompt, the file setup, and the seven-day challenge, ready to print or keep open next to your agent.

Adapted from "I Built a Team of Self-Improving AI Agents" (youtube.com/watch?v=_k5hb5q_yHc). Figures marked as planning estimates are illustrations, not results.