AigentLab® Field Guide / No. 27
Train
The Night
Shift.
10.06.2026
Executive Summary
The Loop That
Grades Itself
Adapted from "I Built a Team of Self-Improving AI Agents". The creator runs agents on a nightly schedule. Each agent makes one piece of content, scores the piece against a written list of criteria, writes down what went wrong, and starts the next night from those notes. He wakes up to a better web page, a better short, and a better draft than the night before.
An agent with four jobs in one cycle. Make something. Grade the result. Write down the lesson. Try again next time with the lesson loaded. The grading and the note-taking are the parts people skip.
The agents run while he sleeps. 1:00 a.m. web pages. 3:00 a.m. short-form videos. 4:00 a.m. Substack. 5:00 a.m. long-form video. Each morning he opens improved systems instead of blank ones. His core line: let agents work while you sleep, not while you work.
In March, Andrej Karpathy published Auto Research, an open-source repo built to let a model improve its own training code overnight. The post drew about 11 million views. The loop repeats four steps all night: change one thing, run a five-minute test, compare the score, keep the change if the score improves and discard the change if the score drops. One night tested about 700 ideas, kept about 20, and finished 11 percent faster than changes made by hand.
You do not need Auto Research or GitHub. You need work you care about and a way to score the work. The same learning loop applies to a web page, a video, or a piece of writing.
Night one produced a clunky, linear page with a scrolling sidebar. Night two scored lower, because the agent had little training to work from. By night three the design, animation, and brand fit beat the best pages he built by hand. The rules the agent writes also govern pages he builds himself.
Prompt, review, complain about colors or text, repeat. The back-and-forth produces good results because you bring the taste. The cost is your time in the chair. The loop hands the critique to an agent trained to judge the way you judge.
He calls the scoring criteria the most important and slowest piece to build. He found them by building pages by hand and noticing what he kept criticizing.
| Area | What he checks |
|---|---|
| Contrast and fonts | Bold, legible type and enough contrast against the background. |
| Sizing | Large elements, with room left for his face in the corner of the video frame. |
| Clarity | No vague "it" or "she" without a named subject. The page reads clearly to a viewer. |
| Visual craft and motion | Glows, frosted panels, animated diagrams. More interest, not always more clarity. |
| Topic fit | One night produced a strong page about pi. Wrong channel, unusable. |
| Layout, type, interaction | Placement of text and graphics, big numbers, no generic serif fonts, clickable elements. |
| Restraint, wow, narrative | Strip what hurts the message, add one striking moment, keep a beginning, middle, and end. |
Each night the agent cycles to one of seven areas and improves only that area. Improving everything at once burns tokens and moves nothing far. When a run goes wrong, the agent writes itself a rule so the mistake does not repeat.
Short-form: a flat stacked graphic became a staircase animation with bounce, glow, and large background numbers, aiming for 10 percent better each night across hooks, scripts, and screenshots. Long-form AI-generated video is next. Substack runs on an agent trained on his Obsidian vault plus universal truths about titles, cover images, and hooks. Every system needs training by hand first.
A markdown file where he leaves dictated notes each morning. The agent reads the feedback first and applies the notes before starting its own loop. He stays in charge without staying in the loop.
The takeaway
Set your taste once, in writing. The agent improves against your written taste every run, and your morning notes steer the next run.New Video Concept / Core Message
Train The
Night Shift.
You fix the same mistakes in every AI draft because nothing you correct carries forward. Write your taste down as a scored rubric. Let an agent make one draft, grade the draft, and log one rule per run, on one focus area. Leave your notes in a file instead of a chat. Every run starts from the last run's rules, so your work improves overnight instead of during your workday.
Opening Premise
You use AI for the work a new professional ships every week: proposals, cover letters, outreach emails, case studies, one-page summaries. Each draft comes back about 70 percent right. You fix the same things every time. A vague opener. A pronoun with no subject. A paragraph too long. A price buried at the bottom. Next week you fix the same things again, because your corrections lived in a chat thread that ended.
The standard advice says write a better prompt. A better prompt helps one draft. The creator of "I Built a Team of Self-Improving AI Agents" hit the same wall with web pages: the back-and-forth worked, and the back-and-forth ate his day. His fix was a loop with a memory, run on a schedule while he slept.
Opening question
Look at your last three AI drafts. How many of your corrections were the same correction?Teaching Points
Keep the loop. Skip the lab.
Karpathy's loop needed code, a GPU, and a five-minute benchmark. Yours needs four moves: make, grade, keep or toss, write the rule.
Pick one asset you produce at least once a week and judge mostly on quality, not facts: a proposal, a cold email, a portfolio case study, a LinkedIn post. Give the agent a fixed input it reuses every run, such as one past client brief. A fixed input means score changes come from the draft, not from a new brief.
| Move | What you set |
|---|---|
| Pick | One asset type, made weekly. |
| Freeze | One input reused every run. |
| Score | The same rubric every run. |
| Compare | Keep the new draft only when the total score rises. |
Rule
One asset. One input. One loop. Add a second loop only after the first runs clean for two weeks.Write your taste before you hand it off.
The creator called the criteria the hardest and most important part. He found his by noticing what he kept criticizing. Do the same with your own edits.
Take your last five drafts. Next to each edit, write why you made the edit in four words or fewer. Group the reasons into five to seven areas. Turn each area into a question scored 1 to 5, with one sentence describing a 5. Write questions a stranger could score. "Sounds good" fails. "Names the client's stated problem in the first two sentences" passes.
Sample rubric: freelance proposal
| Area | Question, scored 1 to 5 |
|---|---|
| Clarity | Does every pronoun point to a named person or thing? |
| Fit | Do the first two sentences name the client's stated problem? |
| Proof | Does the draft include one real number from past work? |
| Length | Is the draft under 400 words with no paragraph over four lines? |
| Restraint | One offer, one price, one next step? |
| Voice | Does the draft sound like you talking across a table? |
| Next step | Does the close name a date and a single action? |
Rule
No rubric, no loop. An agent grading against a vague rubric learns to please itself.One focus area per run.
The creator rotates through seven areas, one per night. Improving everything at once spends tokens and changes little.
Score the whole rubric on every run, so you see whether a fix in one area broke another. Change only the focus area. Keep the new draft when the focus score and the total both rise. Toss the draft when either drops, and log why. Rotate in a fixed order: Monday clarity, Tuesday fit, Wednesday proof, Thursday length, Friday restraint, Saturday voice, Sunday next step.
Rule
Seven areas, seven nights. One area per run, the full rubric scored every time.Turn every miss into a rule.
When a run scores lower, the agent writes one line into a rules file. The next run reads the rules file before writing a word.
Write each rule as an instruction with a reason: "Never open with 'I hope this finds you well.' Clients skip the first line." Specific rules work. "Be more engaging" does nothing. Cap the file at 25 lines. Once a month, merge duplicates, delete rules the drafts no longer break, and resolve any two rules pulling in opposite directions.
Rule
Paste the rules file into your own AI chats too. The night shift trains your day shift.Leave notes, not chats.
The feedback dock is a plain markdown file. Each morning you read the overnight draft, dictate what bothers you into the file, and save.
Date every note and write the fix, not the feeling. "Too long" helps less than "cut the middle section to three bullets." Schedule one run per night at a fixed time, such as 1:00 a.m. ET. A laptop has to be awake and online for a local schedule, so use a cloud scheduled task if your machine sleeps. Expect a dip. The creator's second night scored worse than his first. Judge the loop after seven runs, not one.
Rule
Ten minutes of notes in the morning replaces forty minutes of edits in the afternoon.Translation / The Night Shift, Built
Dana's Proposal Loop
Dana is a first-year freelance copywriter. She sends about six proposals a week to small business owners and spends about 40 minutes editing each AI draft. Every edit repeats last week's edits.
| Piece | Setup |
|---|---|
| Asset | Proposals for small business website copy. |
| Fixed input | One anonymized past brief: a bakery owner wants a new About page and three service pages. |
| rubric.md | Seven questions from her own edit log, 1 to 5 each, 35 points max. |
| rules.md | Empty on night one. 25-line cap. |
| feedback.md | Dated morning notes, dictated in under 10 minutes. |
| outputs/ | One dated draft per run plus a score line. Nothing gets sent. |
| Schedule | Nightly at 1:00 a.m. ET. Focus rotates Monday through Sunday. |
| Boundary | Practice brief only. Live client proposals stay manual until week three. |
Planning estimate: editing time per live proposal drops from about 40 minutes to about 15 by week three, plus 10 minutes of morning notes. Log your own numbers for four weeks and replace these.
The nightly run prompt
Paste this into a scheduled task. Replace the bracketed parts. Keep the order, because the order sets priority: your notes, then the rules, then the work.
- Read feedback.md. Apply every open note first. Mark each note applied.
- Read rules.md. Follow every rule.
- Read rubric.md. Today's focus area: [area for this weekday].
- Read the latest kept draft in outputs/ and the fixed input in inputs/.
- Write one new draft. Change only what improves the focus area.
- Score the new draft and the latest kept draft on all seven questions, 1 to 5, with one sentence of evidence per score.
- Keep the new draft only if the focus score and the total both rise. Otherwise mark the draft tossed.
- If the new draft scored lower anywhere, add one rule to rules.md: one line, one reason. Never exceed 25 lines.
- Save the draft as outputs/[date]-[kept or tossed].md with the score table at the top.
Check yourself
Once a week, score one draft yourself before reading the agent's score. When your total and the agent's total differ by more than 4 points, rewrite the rubric question you disagreed on most.Counterpoint / What the video leaves out
Before You
Trust the Score
The agent scores its own drafts and learns what earns points from itself. Scores rise faster than quality. Spot-check against your own scores every week.
The win arrived on night three of a fresh setup, with no score scale or sample size on screen. Judge your loop over weeks.
Karpathy's loop measured speed with a five-minute test. Taste needs a rubric to stand in for the stopwatch. A vague rubric produces a vague loop.
Four nightly jobs burn tokens and the video never shows a cost. Cap runs at one per night per asset and log spend every Friday.
Every miss adds a line. Without pruning, the agent obeys contradictions. The 25-line cap and monthly cleanup carry the weight.
The pi page looked strong and served nobody. Fix the input and the topic list before blaming the design.
Training on your vault narrows the gap. The draft stays AI-written. Disclose AI use where clients or platforms require disclosure, and send nothing unreviewed.
Closing Challenge / Seven Days, One Loop
Write The Taste.
Then Go To Sleep.
A new professional who used to fix the same five problems in every draft now reads a better draft each morning, leaves ten minutes of notes, and spends the afternoon on clients. Every correction made once stays made.
Pick one asset. Pull your last five drafts and every edit you made.
Tag each edit in four words or fewer. Group the tags into five to seven areas.
Write rubric.md: one question per area, scored 1 to 5, one sentence describing a 5.
Create rules.md and feedback.md. Run the nightly prompt once by hand and score the draft yourself first.
Schedule the run at 1:00 a.m. ET with one focus area per weekday.
Read the overnight draft. Dictate dated notes into feedback.md in ten minutes or less.
Compare scores across runs. Delete one rule. Decide: keep the loop, fix the rubric, or kill the loop.
Train the night shift.
AigentLab® No. 27
Take The Full
Field Guide
15 pages. The full rubric, the nightly run prompt, the file setup, and the seven-day challenge, ready to print or keep open next to your agent.
Adapted from "I Built a Team of Self-Improving AI Agents" (youtube.com/watch?v=_k5hb5q_yHc). Figures marked as planning estimates are illustrations, not results.
