db86edbc21
Capture rendered-text and visual baselines before route migration, and detect external CSS dependencies. Do not change verify.mjs assertions; task 02 must remove the known legacy font URL before the audit can pass.
376 lines
8.4 KiB
Plaintext
376 lines
8.4 KiB
Plaintext
AI For Dummies — Field Guide
|
||
A
|
||
field guide
|
||
01 fleet
|
||
02 worktrees
|
||
03 models
|
||
04 skills
|
||
05 create
|
||
06 field kit
|
||
07 hands-on
|
||
08 verify
|
||
review submissions ↗
|
||
EN
|
||
/
|
||
PT
|
||
AI ENGINEERING
|
||
01 / 2026
|
||
A presentation for humans who ship
|
||
AI for
|
||
dummies.
|
||
You do not need an army of models. You need a system: one mind to frame the work, several hands to execute it, and a clean boundary between every task.
|
||
FIELD NOTE / 001
|
||
Ship the
|
||
system.
|
||
Skills · agents · worktrees · proof
|
||
01
|
||
strong model
|
||
for ambiguity
|
||
03
|
||
bounded workers
|
||
in parallel
|
||
∞
|
||
iterations
|
||
with evidence
|
||
Read this as a route map, not a prompt recipe.
|
||
RULE ZERO
|
||
Strong model for ambiguity.
|
||
Light model for bounded work.
|
||
THINK
|
||
MAKE
|
||
A small fleet
|
||
coordination before parallelism
|
||
ORCHESTRATOR
|
||
Decides what
|
||
needs to happen.
|
||
Opus / reasoning
|
||
→
|
||
UI
|
||
Component and visual states
|
||
agent/ui
|
||
TEST
|
||
Acceptance cases
|
||
agent/tests
|
||
DOCS
|
||
Guide and examples
|
||
agent/docs
|
||
The orchestrator preserves intent, writes small contracts, and gathers results that can be verified. It does not need to type every line.
|
||
Why the boundary matters
|
||
one vague task / three predictable failures
|
||
01
|
||
Context soup
|
||
Every worker reads everything. Nobody knows which facts are load-bearing.
|
||
02
|
||
Branch collision
|
||
Two agents touch the same checkout. The fastest path becomes conflict resolution.
|
||
03
|
||
Confident drift
|
||
The diff is polished, but no one checks whether it solved the original problem.
|
||
The subagent loop
|
||
Click a phase.
|
||
See the handoff.
|
||
Delegation means moving one bounded task into a smaller context—not giving away responsibility.
|
||
01
|
||
PLAN
|
||
02
|
||
BUILD
|
||
03
|
||
REVIEW
|
||
What crosses contexts
|
||
brief → diff → evidence
|
||
Package
|
||
Contains
|
||
Why it matters
|
||
Brief
|
||
goal, files, boundaries
|
||
stops the worker inventing the problem
|
||
Worktree
|
||
branch and isolated checkout
|
||
parallel edits do not collide
|
||
Checks
|
||
tests, build, criteria
|
||
turns “looks good” into evidence
|
||
Diff
|
||
small, reviewable change
|
||
integration and discard stay cheap
|
||
Git worktrees
|
||
One branch
|
||
per
|
||
hand.
|
||
A worktree is another directory linked to the same repository. Each agent gets its own checkout and index; history remains shared.
|
||
Select a node to inspect its checkout, owner, and next action.
|
||
repository topology
|
||
4 checkouts
|
||
ROOT
|
||
main
|
||
● clean
|
||
UI AGENT
|
||
agent/ui
|
||
3 files · working
|
||
TEST AGENT
|
||
agent/tests
|
||
8 checks · ready
|
||
DOCS AGENT
|
||
agent/docs
|
||
2 pages · review
|
||
Model routing
|
||
Do not pay for
|
||
reasoning where
|
||
you need
|
||
rhythm.
|
||
Choose a job to see why the model profile changes.
|
||
Work
|
||
Profile
|
||
Prompt shape
|
||
Plan
|
||
strong / broad
|
||
What changes? What can break?
|
||
Build
|
||
fast / focused
|
||
Implement this slice. Run these checks.
|
||
Explore
|
||
read-only / light
|
||
Find where this contract is used.
|
||
Review
|
||
independent
|
||
Does the diff satisfy the brief?
|
||
Model gearbox
|
||
capability tier × thinking effort
|
||
Two separate knobs
|
||
Choose the engine.
|
||
Then choose the
|
||
gear.
|
||
A stronger model changes the capability ceiling. Higher reasoning effort gives that model more room to work. Start with the lightest combination that passes your real checks, then move one knob at a time.
|
||
OPENAI
|
||
CLAUDE
|
||
GEMINI
|
||
REASONING / THINKING
|
||
LOW
|
||
bounded + fast
|
||
MEDIUM
|
||
default start
|
||
HIGH
|
||
complex + costly
|
||
ROUTING RULE
|
||
Use strong models for ambiguity and judgment. Use lighter models for bounded execution. Raise effort only when evaluation shows a gain.
|
||
Skills
|
||
Write the right way
|
||
once.
|
||
A skill is a reusable procedure. It can carry instructions, references, scripts, and assets. It is not magical memory, and it does not replace acceptance criteria.
|
||
01 / trigger clearly
|
||
02 / load detail on demand
|
||
03 / return evidence
|
||
SKILL PACKAGE
|
||
SKILL.md
|
||
procedure and limits
|
||
references/
|
||
facts to consult
|
||
scripts/
|
||
repeatable checks
|
||
assets/
|
||
templates and examples
|
||
name: review-ui · check focus, mobile, reduced motion · run verification · return evidence
|
||
Create a skill
|
||
repeatable pain → reusable judgment
|
||
The skill forge
|
||
Teach the decision.
|
||
Keep the context
|
||
light.
|
||
Do not package everything you know. Capture the non-obvious choices that repeatedly improve an outcome, then prove the skill changes behavior.
|
||
01
|
||
Observe
|
||
find repeated friction
|
||
02
|
||
Define trigger
|
||
route precisely
|
||
03
|
||
Choose anatomy
|
||
only needed files
|
||
04
|
||
Write guidance
|
||
decisions, not trivia
|
||
05
|
||
Validate
|
||
test real behavior
|
||
OUTPUT / SKILL PACKAGE
|
||
review-ui/
|
||
├── SKILL.md
|
||
├── agents/
|
||
│ └── openai.yaml
|
||
├── references/
|
||
│ └── accessibility.md
|
||
└── scripts/
|
||
└── verify.mjs
|
||
VALIDATE
|
||
quick_validate.py ./review-ui
|
||
AFTER REAL USE
|
||
observe failure
|
||
→
|
||
sharpen one rule
|
||
→
|
||
retest behavior
|
||
→
|
||
keep it narrow
|
||
Common skills
|
||
choose behavior before model
|
||
The field kit
|
||
Different jobs.
|
||
Different
|
||
instincts.
|
||
A skill changes how an agent approaches work. Some shape communication. Others enforce research, debugging, review, or completion discipline. Select one to inspect its operating rule and verified source.
|
||
SIMPLIFY
|
||
ponytail-lite
|
||
minimum code that holds
|
||
COMMUNICATE
|
||
caveman
|
||
signal without filler
|
||
COMPLETE
|
||
unlazy
|
||
gates and evidence
|
||
INVESTIGATE
|
||
research
|
||
primary sources first
|
||
DIAGNOSE
|
||
diagnosing-bugs
|
||
tight feedback loop
|
||
REVIEW
|
||
code-review
|
||
standards × spec
|
||
ECONOMIZE
|
||
token-saver
|
||
compress noisy output
|
||
ONE PRACTICAL LOADOUT
|
||
PLAN
|
||
unlazy
|
||
→
|
||
BUILD
|
||
ponytail-lite
|
||
→
|
||
DEBUG
|
||
diagnosing-bugs
|
||
→
|
||
REPORT
|
||
caveman
|
||
INSTALL PACK
|
||
Ask your coding agent to verify, install, and validate the skills.
|
||
COPY
|
||
↗
|
||
Review every source before installation. Existing local skills must be preserved.
|
||
Hands-on
|
||
10 minutes / one missing feature
|
||
Tiny Tasks lab
|
||
Same task.
|
||
Better
|
||
operating system.
|
||
Start with a deliberately incomplete static task board. Run one prompt as written, reset, then run the skill-enabled version. Compare diff size, verification evidence, and unnecessary complexity.
|
||
Open the starter →
|
||
Clone from Gitea →
|
||
Open the rules lab →
|
||
Clone from Gitea →
|
||
THE MISSING FEATURE
|
||
Add All / Open / Done filters that survive reload and browser navigation.
|
||
STACK
|
||
HTML · CSS · JavaScript
|
||
DEPENDENCIES
|
||
none
|
||
FILES
|
||
3
|
||
RUN A
|
||
Good prompt
|
||
COPY
|
||
↗
|
||
Clear context · constraints · acceptance · evidence
|
||
RUN B
|
||
Good prompt + skills
|
||
COPY
|
||
↗
|
||
Same contract · explicit working methods · stronger proof
|
||
COMPARE THE RUNS
|
||
01
|
||
Files changed
|
||
02
|
||
New dependencies
|
||
03
|
||
Checks actually run
|
||
04
|
||
Evidence returned
|
||
THE HUMAN JOB
|
||
The agent may be autonomous in execution. Intent, boundaries, and evidence remain yours.
|
||
START HERE
|
||
Begin with one agent and one skill. Add parallelism only when the tasks are truly independent.
|
||
Verification
|
||
run each gate separately
|
||
Checks become evidence
|
||
Three layers.
|
||
Run each one alone.
|
||
Run a gate on its own line, print its exit code, attach the output. The result is the deliverable.
|
||
01 · STATIC
|
||
Lint and types
|
||
Format, lint, type-check. Fast and scoped to one file. Run on every save.
|
||
pnpm lint; echo "lint=$?"
|
||
pnpm typecheck; echo "typecheck=$?"
|
||
02 · BEHAVIOR
|
||
Unit and contract
|
||
Tests that repeat. Run before claiming done.
|
||
pnpm test; echo "test=$?"
|
||
cd services/api && go test ./...
|
||
03 · INTEGRATION
|
||
Real UI and API
|
||
Drive the actual UI, API, or browser. Slower and flakier — only this catches mobile overflow and a missing 404.
|
||
pnpm check:ui; echo "ui=$?"
|
||
TURBO_FORCE=true pnpm e2e
|
||
FOUR WAYS A GREEN REPORT IS FALSE
|
||
1
|
||
Pipe a gate
|
||
tail, grep, or head hide the real exit code — a pipeline returns the last command's status.
|
||
2
|
||
Swallow a rejection
|
||
A silent
|
||
.catch(() => {})
|
||
hides a panic, an upstream limit, or a partial failure.
|
||
3
|
||
Trust the cache
|
||
Turbo caches results. A gate that "passes" may not have run — use
|
||
TURBO_FORCE=true
|
||
.
|
||
4
|
||
Skip the third layer
|
||
Lint and unit can both be green while the page breaks on mobile and the API never returns 404.
|
||
RUN IT YOURSELF · two labs, under 10 minutes each
|
||
Path A · verification lab
|
||
Fill the four-row comparison strip on the starter. Run A naively, Run B with
|
||
$gate-discipline
|
||
and
|
||
$webapp-testing
|
||
.
|
||
Open the starter →
|
||
Clone ↗
|
||
git.marcospaulo.dev.br/.../src/branch/pages/hands-on/starter
|
||
Path B · rules lab
|
||
Toggle every rule off, run the prompt. Toggle every rule on, run it again. Compare diff size, gate invocations, and the names of checks the agent names back.
|
||
Open the rules lab →
|
||
Clone ↗
|
||
git.marcospaulo.dev.br/.../src/branch/pages/hands-on/rules
|
||
Keep learning
|
||
12 new readings + primary docs
|
||
Go deeper with official documentation, production case studies, Medium, and practitioner workflows.
|
||
Rules and enforcement case study →
|
||
Skills review desk →
|
||
Primary references →
|
||
12-part reading path →
|
||
Navigate by idea
|
||
short chapters / one system
|
||
Prefer a focused chapter? Start with the
|
||
route map
|
||
, then jump directly to
|
||
models
|
||
,
|
||
agents and worktrees
|
||
,
|
||
skill creation
|
||
,
|
||
rules
|
||
, or the
|
||
skills review desk
|
||
.
|