To test local web applications, write native Python Playwright scripts.
Helper Scripts Available:
scripts/with_server.py- Manages server lifecycle (supports multiple servers)
Always run scripts with --help first to see usage. DO NOT read the source until you try running the script first and find that a customized solution is abslutely necessary. These scripts can be very large and thus pollute your context window. They exist to be called directly as black-box scripts rather than ingested into your context window.
Decision Tree: Choosing Your Approach
User task → Is it static HTML?
├─ Yes → Read HTML file directly to identify selectors
│ ├─ Success → Write Playwright script using selectors
│ └─ Fails/Incomplete → Treat as dynamic (below)
│
└─ No (dynamic webapp) → Is the server already running?
├─ No → Run: python scripts/with_server.py --help
│ Then use the helper + write simplified Playwright script
│
└─ Yes → Reconnaissance-then-action:
1. Navigate and wait for networkidle
2. Take screenshot or inspect DOM
3. Identify selectors from rendered state
4. Execute actions with discovered selectors
Example: Using with_server.py
To start a server, run --help first, then use the helper:
Single server:
python scripts/with_server.py --server "npm run dev" --port 5173 -- python your_automation.py
Multiple servers (e.g., backend + frontend):
python scripts/with_server.py \
--server "cd backend && python server.py" --port 3000 \
--server "cd frontend && npm run dev" --port 5173 \
-- python your_automation.py
To create an automation script, include only Playwright logic (servers are managed automatically):
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True) # Always launch chromium in headless mode
page = browser.new_page()
page.goto('http://localhost:5173') # Server already running and ready
page.wait_for_load_state('networkidle') # CRITICAL: Wait for JS to execute
# ... your automation logic
browser.close()
Reconnaissance-Then-Action Pattern
Inspect rendered DOM:
page.screenshot(path='/tmp/inspect.png', full_page=True) content = page.content() page.locator('button').all()Identify selectors from inspection results
Execute actions using discovered selectors
Common Pitfall
❌ Don't inspect the DOM before waiting for networkidle on dynamic apps
✅ Do wait for page.wait_for_load_state('networkidle') before inspection
Best Practices
- Use bundled scripts as black boxes - To accomplish a task, consider whether one of the scripts available in
scripts/can help. These scripts handle common, complex workflows reliably without cluttering the context window. Use--helpto see usage, then invoke directly. - Use
sync_playwright()for synchronous scripts - Always close the browser when done
- Use descriptive selectors:
text=,role=, CSS selectors, or IDs - Add appropriate waits:
page.wait_for_selector()orpage.wait_for_timeout()
Always Measure Above-the-Fold, Not Page Completion
Run this on every web app you touch, including B2B and internal admin tools. The instinct is that consumer apps need to be fast and internal dashboards can be slow because "it's just the team". That is backwards in one important way: internal tools are used all day, every day, by people who cannot choose a competitor, and a 10-second dashboard silently costs more hours than a slow marketing page ever will.
The metric trap
Do not report "the app takes N seconds" from a curl/fetch timing. Reading a response to
the last byte measures PAGE COMPLETION. On a streamed page (React Suspense, RSC, Turbo frames,
htmx out-of-band swaps), completion is dominated by whatever below-the-fold section is slowest
— the exact thing streaming exists to decouple from the user's wait.
Measured on a real Next.js admin dashboard, same route, same moment:
| what was measured | number |
|---|---|
fetch() to last byte, summed over 34 routes |
29.7s |
| TTFB, worst route | 615ms |
| LCP, worst route | 1,608ms |
Both are true. Only the second pair describes what a human experiences. Summing completion times across routes is close to meaningless — nobody loads 34 pages at once.
Above the fold is TTFB, FCP, LCP and CLS. Everything else is a secondary diagnostic.
Measure it properly
performance.getEntriesByType("largest-contentful-paint") after load returns NOTHING. LCP and
CLS are only delivered through a PerformanceObserver, and the observer has to exist before
the paint. Install it with addInitScript before navigating, or you will silently record
null for the one metric that matters. Use scripts/measure_above_fold.py.
Report per route, and quote the WORST route, never the average:
route TTFB FCP LCP CLS
/dashboard 148ms 176ms 176ms 0
/reports 208ms 240ms 644ms 0
Thresholds (Core Web Vitals): LCP <= 2.5s good, <= 4s needs improvement. CLS <= 0.1. TTFB <= 800ms.
Cold and warm are different apps. Measure both.
Warm numbers flatter you and hide the bug. The same dashboard:
| route | warm LCP | cold LCP |
|---|---|---|
/analytics |
644ms | 900ms |
/taxonomy |
776ms | 18,076ms |
/taxonomy had a 228ms TTFB in both. The shell streamed instantly while the content the user
actually came for was 18 seconds away. A fast TTFB on a streamed page can hide an unusable
page — that is why LCP is the metric and TTFB alone is a trap.
To force cold: clear the framework data cache (rm -rf .next/cache for Next.js), restart the
server, and take the FIRST hit on each route. Then ask which of these a real user hits:
- after a deploy: everyone
- after a cache TTL expires: one user per key per TTL
- after a scheduled job rewrites the data: the first user of every cycle
That last one is the one people forget, and it recurs forever.
Precompute above-the-fold data. Especially in B2B.
If a number renders above the fold, it should be READ, not COMPUTED, at request time. A dashboard stat that needs a 12-CTE aggregate over an event log does not become acceptable because the audience is internal.
The ladder, cheapest first:
- Cache the read (
unstable_cache, Redis, HTTP cache). Fixes repeat views, not the first. - Precompute on write (materialized view, rollup table, counter column). Fixes every view.
- Warm the cache as part of the job that invalidates it. The step people skip. If a cron refreshes a materialized view, every page of it is cold immediately afterwards and the first human of the hour eats the disk reads. The job is already awake and already paying — have it run the two or three heaviest read shapes before it exits. Costs seconds of a machine's time to save seconds of a person's, every cycle.
- Stream the rest. Put a Suspense boundary around anything that is genuinely slow AND genuinely below the fold. Never around the primary content — that just moves a blank screen behind a spinner.
Anti-pattern to check for specifically: a page that fetches every tab's data before rendering the first tab. Above the fold only needs the default view.
Checklist
- TTFB, FCP, LCP, CLS captured per route with a
PerformanceObserver, not afetchtimer - Cold AND warm measured; cold taken after clearing the data cache
- Worst route quoted, not the average
- Any LCP > 2.5s traced to a specific query or payload, not hand-waved
- Above-the-fold data precomputed, and the precompute job warms what it invalidated
- Suspense/streaming used only below the fold
- DOM node count checked (>5,000 on one screen usually means the page renders rows nobody has scrolled to)
Reference Files
- scripts/measure_above_fold.py - TTFB/FCP/LCP/CLS per route, cold and warm
- examples/ - Examples showing common patterns:
element_discovery.py- Discovering buttons, links, and inputs on a pagestatic_html_automation.py- Using file:// URLs for local HTMLconsole_logging.py- Capturing console logs during automation