Review model errors by hand, discover failing slices, and decide whether data or model fixes pay more.
Write role cards that pin each agent to one charter with declared inputs, outputs, refusal rules, and non-overlapping scope.
Precompute expensive read queries into materialized views and pick a refresh strategy that fits a stated staleness budget.
Design multi-turn conversational systems with state tracking, memory injection, topic handling, and repair.
Store, compare, and bucket timestamps with explicit time zones so reports do not drift by a day and ranges do not miss rows.
Force the thread interleavings that expose races by using barriers, stress loops, and linearizability checks instead of hoping the scheduler hits them.
Control the clock in tests through an injectable time source so "now" is frozen and timezone behavior is explicit.
Back-of-envelope calculations for system design. Use when estimating QPS, storage, bandwidth, or latency for capacity planning.
Read EXPLAIN ANALYZE output to find the real cause of a slow query and fix the right thing. Use when a query is slow and you need to know why before changing indexes or SQL.
Batch many low-urgency events into one periodic summary that is worth opening, rather than sending each as it happens.
Operate as a principal architect who owns the boundaries between systems, curates the technology set, and spends veto power sparingly.
Cascade OKRs so team and individual goals ladder up to company strategy without sandbagged targets or scoring that rewards easy wins.
Instrument code with counters, gauges, and histograms picked to answer a specific operational question.
Match words to their variants and equivalents without collapsing distinctions that matter. Use when searches miss obvious results, or when unrelated results appear because two…
Run a data team as agents that build the pipeline, gate on quality checks, run the analysis, and independently audit every headline metric.
Write formulas that stay correct when rows are added and readable when someone else opens the file. Use when building any calculation in a spreadsheet.
Design experiments with controls, randomization, confound awareness, and pre-registered analysis. Use when testing a hypothesis empirically and needing the result to actually mean…
Reduce input/output cost by cutting round trips, moving fewer bytes, and turning random access into sequential access.
Apply complexity analysis where input size actually makes it decide performance, and ignore it where constant factors dominate.
Prepare a raise with agents that assemble the narrative, stress-test it as an investor would, and organise diligence material, while the founder owns every conversation.
Get data into a search index reliably and keep it current, with reindexing, partial updates, and a defined staleness budget.
Keep git history atomic, bisectable, and readable with curated commits and disciplined force-push. Use when commits are messy, history is hard to navigate, or bisect and blame…
Diagnose and fix slow spreadsheets by reducing volatile formulas, whole-column references, and unnecessary recalculation. Use when a workbook takes seconds to respond to an edit.
Replace unexplained literal values with named constants that carry their meaning, without over-formalizing the obvious. Use when a bare number or string encodes a rule.
Write bash that fails loudly, quotes correctly, cleans up after itself, and passes shellcheck. Use when writing shell scripts that automation or other people will depend on.
Express limits on length, scope, style, and behaviour so they are followed rather than politely ignored. Use when a model consistently exceeds bounds or drifts outside the task.
Write the durable instructions that shape every turn, covering role, boundaries, and defaults without micromanaging each response.
Compose specialist agent desks into one operating company with a shared ledger, a weekly cadence, and human approval gates on anything binding.
Replicate a live incident team as agents with a commander, parallel investigators, comms, and a scribe, coordinated on a fixed cadence.
Turn support conversations into product change through structured tagging, aggregation, and a route into the roadmap.
Build Python command-line tools with correct exit codes, stream discipline, and a distributable entry point. Use when writing or packaging a CLI in Python.
Run a team of support agents that triages each ticket, reproduces the problem, drafts a reply, and hands clean escalations to a human.
Organise a repository so a newcomer finds what they need and automation has predictable paths. Use when starting a repository or when nobody can find anything in an existing one.
End a project deliberately, with handover, documentation, and a decision about what happens to what was built. Use when a project reaches its goal or is being stopped.
Comprehensive guide to load balancing algorithms, health checks, and high-availability patterns.
Find the object that never gets freed by comparing heap snapshots and following retention paths. Use when a process grows in memory over time, gets OOM-killed, or slows under a…
Convert uploaded video into the renditions needed for reliable playback across devices and bandwidths. Use when accepting user video or delivering video that must play everywhere.
Isolate render failures behind boundaries with useful fallbacks, retry, and reporting so one broken component does not blank the app.
Tag releases with semantic versions and signed tags, generate changelogs, and trigger builds from tags.
Define what a user sees when a string, a locale, or a region is not available, so gaps degrade predictably instead of showing keys or blanks.
Write Makefiles with honest dependencies, phony discipline, and self-documenting help, and know when make is the wrong tool.
Choose batch or streaming from honest latency requirements and operate the complexity you actually need.
Reduce overfitting with weight decay, dropout, augmentation, and early stopping, choosing by why the model is overfitting.
Build a reliable, minimal reproduction before you write a fix so you can prove the bug is actually gone.
Add redundancy where it removes a single point of failure, understanding what each level protects against and what it costs. Use when designing for availability targets.
Use CTEs to name intermediate steps so a complex query reads as a sequence rather than a nest, and know when they cost performance.
Run a lean startup team of agents (builder, marketer, analyst) through one disciplined weekly loop of ship, promote, and measure.
Record decisions with their context and predictions to enable honest calibration and defeat hindsight bias. Use when making consequential decisions you want to learn from later.
Understand what Raft and Paxos actually provide, quorum arithmetic, and when you need consensus at all.
Detect and reverse the point where users stop reading, by measuring engagement decay and cutting volume before they mute. Use when open rates are falling or mute rates are rising.
Choose an optimiser and its hyperparameters based on the problem rather than habit, and know what each actually does.
Assemble a promotion packet that proves sustained impact at the next level, with calibrated scope claims and evidence a committee can verify.
Size worker pools, queue depths, and batch widths against the real bottleneck so parallelism adds throughput instead of contention.
Let teams provision what they need within guardrails, without waiting for an infrastructure team to act. Use when provisioning requests queue and infrastructure work is reactive.
Suggest queries as the user types, fast enough to feel instant and relevant enough to be worth reading.
Give a coding agent the right files and background at the start of a task, so it works from the real system rather than from assumption.
Script repository operations reliably, respecting rate limits, pagination, and permissions. Use when managing repositories at scale or building tooling around the platform.
Know whether a message was accepted, delivered, opened, and acted on, and treat failures as events rather than statistics.
Price SaaS around a value metric with tier design, seat-vs-usage decisions, and low-risk price testing. Use when setting or revisiting software pricing and packaging.
Improve ranking with a judged evaluation set and measured changes rather than intuition. Use when search feels wrong and every proposed fix is someone's opinion.