Jigon YooAI & data reliability Email me

Jigon Yoo · Freelance AI & data reliability engineer

I find where your AI agents and data pipelines go wrong, and prove the fix with numbers.

You get a reproducible test that shows the failure, a written report on why it happens, and a before-and-after measurement once it is fixed. Fixed scope, fixed price, all in writing.

A finding, measured 96 runs each

How often does an AI support agent refund the same customer twice?

Paid twice: refunds paid twice in 96 runs per model, under timeouts, re-filed tickets and parallel workers. $ extra: the total paid twice or over the refund cap across those 96 runs. Scored from a hash-chained ledger, never from what the model says. See the test bench ↗

Services

Three fixed-scope diagnostics

Each one starts from a specific question about your system and ends with an answer you can rerun. Send a short description and I reply with a fixed-price quote and the exact scope.

“Is the new model or prompt actually better?”

Model & prompt switch check

Before you change models or rewrite a prompt, see what changes on your real tasks.

You send
Your current prompt or workflow and representative cases. If you have none, we write a test set together.
You get
A fixed test set, a script you can rerun, and a report comparing success rate, latency and cost before and after.

Same method as the three-model refund-desk measurement above, and the CI gate in eval-harness that catches a 100% → 48% drop.

From $500 usually 5 business daysAsk

“Can this charge or refund happen twice?”

Duplicate side-effect diagnostic

For AI agents and webhook handlers that move money, send messages or write records, under retries, timeouts and parallel workers.

You send
Access to the code that handles tool calls or webhooks.
You get
A failing test that reproduces the duplicate, a report on the cause, a proposed fix as a pull request, and proof the test passes after it.

Evidence: duplicate-side-effect-desk, and agent-reliability-kit, where double charges go from 5 to 0. Write-up: why checking the ledger first still paid twice.

Free tool Try replay-twice on your own handler first: it delivers the same event four ways and checks your ledger counts one effect.

From $800 usually 5 business daysAsk

“Did that batch load correctly, or just without errors?”

Data load quality gate

For warehouses and pipelines where a bad batch loads without a single error and quietly skews every dashboard downstream.

You send
A sample of your data and your schema or dbt project.
You get
A data contract written as dbt tests, a report on faults planted on purpose and how many it caught, and the gate wired into CI.

Evidence: in warehouse-quality-gate a sabotaged batch loads with zero errors while one row carries $4.5M; the contract fails 12 of 15 tests and the reporting tables are not rebuilt. Also sheet-to-db-contract: 31 findings on an import that loaded every row. Write-ups: what unique + not_null let through and the batch that added $4.5M.

Free tool Try dbt-fault-drill on your own dbt project first: it plants up to 17 realistic faults and counts how many your tests stop.

From $600 usually 5 business daysAsk

Prices are starting points for the smallest scope. Your quote fixes the final price and scope in writing before any work starts. Model API usage runs on your own account.

Also on request

Five more fixed-scope checks. Tap one to see what it covers.

  1. Step 1You email meWhat the system does and what worries you.
  2. Step 2Fixed-price quoteI reply with questions or a quote and the exact scope.
  3. Step 3I reproduce and measureEvery finding comes with a test you can run.
  4. Step 4You get the reportFindings, fixes and before-and-after numbers.
Email me

Describe the problem in a few lines. I reply with questions or a fixed-price quote.

Results

Numbers you can rerun yourself

All from public repositories with planted faults and synthetic or sample data. Clone one, run the command in its README, and you should see the same number.

0 errors

A sabotaged batch loads into the warehouse without a single error. One row carries $4.5M. A dbt contract fails 12 of 15 tests and stops the rebuild.

warehouse-quality-gate ↗

5→0

Double charges when the same workload runs with retry only, then with idempotency, timeouts and a circuit breaker.

agent-reliability-kit ↗

100%→48%

Row-level accuracy after a real amount-parsing bug slips into an extraction pipeline. The CI gate blocks the release.

eval-harness ↗

33 of 33

Prompt-injection attacks and secret leaks stopped by a guardrail layer, with 0 false positives on 17 harmless lookalikes.

llm-guardrails ↗

10 of 24

Bad batches a lazy baseline loads while still scoring 0.875, because it never checks the warehouse. The README reports it instead of hiding it.

bad-batch-gate ↗

31 findings

On a spreadsheet import that loaded every row without an error: #REF! stored as text, a missing column, an id turned into 1.23457E+11.

sheet-to-db-contract ↗

Work

The projects behind each service

Start with the nine closest to the services above. Each has a README and a run you can reproduce. All 79 public projects are one click away.

Featured projects

79 projects

Agent safety & tool calls

12 projects

For teams putting LLM agents in front of real tools. Hire me to check what your agent is allowed to do, what it does twice, and what the logs can prove afterwards.

replay-twice

A pytest plugin that delivers the same webhook event or agent tool call again — twice in a row, after a lost reply, after a lost acknowledgement, and eight at once — and asks your own ledger whether the refund happened exactly once. Sync and async def handlers; offline, no model or API key.

Why it matters. A handler that applies the effect once but raises on every redelivery is flagged too, because the sender keeps retrying. On the bundled examples, check-then-act blocks sequential duplicates yet double-applies in 200 of 200 concurrent trials (its race is widened on purpose); the guarded handler stays at 0. Reviewed independently twice; every finding fixed (see its STATUS_REPORT).

  • pytest
  • idempotency
  • webhooks
  • ai agents

agent-approval-gate

An LLM support agent that cannot spend money without a human — and the number that proves it. The same 37 recorded incidents run twice, once through a typical agent loop and once through seven gates: unapproved money moved $453.50 → $0.00, duplicate side effects 4 → 0, injected args executed 4 → 0, malformed money moves 2 → 0, and 5 actions held for a person.

Why it matters. It measures the right thing. Most agent demos score whether the agent finished; this one scores what it moved, replaying every refund that actually landed against policy.yaml to ask whether a human was required. Noise on control tasks stays at 0 — legitimate work is untouched.

  • agent safety
  • human-in-the-loop
  • approval gate
  • deterministic

mcp-permission-server

A permission layer in front of MCP tools, and a log that can prove what it decided. The same 17-call session through a naive allowlist and through nine checks: calls executed without a grant 7 → 0, decisions that cannot be reconstructed from the log 16 → 0, denials that left no trace 1 → 0, secrets sitting in the log 2 → 0, entries removable unnoticed 1 → 0 — with 0 legitimate calls refused.

Why it matters. The log is the point. Almost every audit trail records what happened; almost none record enough to recompute why it was allowed. --verify re-decides all 17 entries from the log alone and checks each against the verdict it claims. Redaction at write time, hash-chained entries, and a stated limit: this is a policy and audit layer, not a sandbox.

  • mcp
  • authorization
  • audit-log
  • hash-chain

agent-reliability-kit

The reliability layer an agent’s tool calls need — retry, timeout, circuit breaker, idempotency — with the numbers showing what each buys. Same workload through naive / retry-only / reliable: naive 53/80 → reliable 63/80, double-charges 5 → 0, load on a down service 88 → 30 calls.

Why it matters. It shows the costs, not just the wins: retry-only’s double-charge trap (5 duplicate side effects) and the breaker’s ~1-request post-recovery cost. Every run is reproducible byte for byte. Offline, no keys.

  • agentops
  • reliability
  • circuit-breaker
  • idempotency

llm-guardrails

A drop-in input/output guardrail layer for LLM apps: blocks prompt injection and jailbreaks before they reach the model, and redacts leaked secrets/PII before they reach the user. Naive lets 33/33 attacks+leaks through; guarded stops all 33 — with zero false positives on 17 benign lookalikes.

Why it matters. The numbers are honest: weighted multi-signal detection that knows use from mention (0 false positives), a Luhn-checked redactor that never mangles clean text, and a CI gate that fails if protection regresses. Offline, no keys.

  • llm security
  • prompt-injection
  • owasp-llm
  • pii-redaction

mcp-server-prod

A hardened MCP server exposing read-only order tools to an AI agent — treating every tool argument as untrusted, because it's LLM-generated. Validated, parameterized, read-only. A naive concat server leaks the whole table on a classic OR-1=1 injection; this one blocks 5/5 with zero functionality lost.

Why it matters. It covers the security most MCP demos skip: tool args are a network boundary, so validation + parameterized SQL + a read-only connection — proven by a contract suite that also shows the naive version really leaks.

  • mcp
  • security
  • sql-injection
  • sqlite

odds-mcp

An MCP server where every read names its evidence and the one write is impossible without an approval token. A risk envelope checked before a human ever sees the request, a single-use token issued by a separate step, a publish that cannot happen without it, and an append-only log of every attempt — especially the refused ones.

Why it matters. The risky tool is built safely. Reads are easy to make safe; this is the propose → approve → publish pattern that makes a side-effectful agent tool safe enough to exist at all. Data comes from odds-consensus, so the integration surface is a real pipeline rather than a mock.

  • mcp
  • approval token
  • audit log
  • agent tools

readonly-guard

A narrow Python guard for audit, inventory, and verification tools that are meant to read evidence without changing the environment they are pointed at. Blocks write-capable open modes, selected pathlib/os/shutil mutations, and HTTP POST/PUT/PATCH/DELETE — while the sample inventory still runs inside the guard.

Why it matters. Its limits are part of the contract: this is a tested application-level policy, not an operating-system sandbox, and the one-pager says so before a client ever sees it.

  • read-only
  • audit tooling
  • safety
  • unittest

agentops-tool-orchestrator

A single-agent loop — planner → guardrail validation → executor → eval harness. Tool registry with arg schemas, allow/deny + approval gates for side-effecting tools, step caps, and loop detection.

Why it matters. The safety + evaluation scaffolding around an LLM agent, not a model: side-effecting tools never run without an approval gate, and every block logs its rule.

  • agentops
  • guardrails
  • evaluation

agentops-multiagent-supervisor

A supervisor decomposes a task and routes subtasks to specialized worker agents over a message bus + shared blackboard, with coordination guardrails.

Why it matters. Multi-agent coordination plumbing with cycle/deadlock detection, delegation-depth caps, and per-agent budgets — deterministic and observable, not emergent magic.

  • agentops
  • multi-agent
  • orchestration

agentops-trace-cost-guard

Wraps an agent run to produce a structured trace (steps, tokens, latency), cost accounting from a pricing table, hard token/$ budget guards, and failure policies (retry, circuit-breaker, fallback).

Why it matters. Makes agent runs observable, budgeted, and fail-safe; costs from a static illustrative table and a simulated clock keep every run reproducible.

  • agentops
  • observability
  • reliability

agent-trace-triage

Agent execution traces → a rule-based failure taxonomy (tool-selection error, loop, timeout, malformed args, wrong answer…), a failure distribution, and a labeled eval-ready dataset with a human-review queue for ambiguous cases.

Why it matters. The missing link between observability (89% adoption) and evaluation (52%): it turns raw traces into eval data deterministically, and sends ambiguous cases to a human rather than guessing.

  • agentops
  • evaluation
  • observability

Evaluation, graders & LLM ops

7 projects

For teams that need to know whether a model, prompt or grader is telling the truth. Hire me to build the test set, attack the grader, or measure a model switch.

duplicate-side-effect-desk

Does the agent make the same real-world thing happen twice? A refund desk where the payment tool sometimes reports a timeout for a payment that went through, and the ledger does not show a write the instant it lands. The adversary is not a jailbreak — it is a retry, a ticket filed twice, and a second worker halfway through the same refund. Three models over 96 rollouts each: duplicate payments 14 / 3 / 3 and $708 / $150 / $205 paid twice or over the cap.

Why it matters. Nothing the model says is scored — every number is recomputed from a hash-chained ledger. Six adversarial agents ship with it and the suite fails if any of them scores like a careful one; three grader holes were closed that way, including one where an API outage raised the score because a run that never happened caused no damage. Telling the model the rule in its prompt bought nothing: it quoted the rule and broke it in the next tool call.

  • rl environment
  • idempotency
  • audit-log
  • verifiers

bad-batch-gate

Does the agent stop a bad batch before it loads? A nightly warehouse load where the loader always answers errors: 0 — a blank amount becomes 0, a re-sent file loads twice, a re-added customer doubles their revenue in every join. 32 eval cases, 24 defective batches among 64. Loading everything leaves the revenue mart $746,410 off across the 32 cases; refusing everything leaves it $567,192 short and scores lower still (0.3875 vs 0.5344).

Why it matters. Scored from what reached the warehouse and a hash-chained log, never from what the agent says. Eight attacker agents ship with it, three of them handed the answers; all but one stay at least 0.05 below the reference — the exception is a realistic mistake, load order, at 0.975. A lazy baseline that profiles every batch but never checks the warehouse scores 0.875 while loading 10 of the 24 bad batches. Independent reviews found the holes, from a quarantine filed after the load to batch ids that leaked the answer; each fix has a test that fails when the bug is put back. Not yet run on real models — the README says what a perfect score takes and why models may crowd near it.

  • rl environment
  • data quality
  • sql
  • verifiers

eval-harness

Golden-set regression testing for LLM / extraction pipelines: scores a pipeline, judges the fields where exact match is unfair, and fails CI when quality regresses. A real amount-parsing bug drops row-correct 100% → 48% and the gate blocks it.

Why it matters. The value is not the metric but a CI gate that catches silent regressions, an LLM-as-judge that recovers correct-but-non-exact answers, and a judge that's itself checked against human labels. Offline, no keys.

  • eval
  • llm-as-judge
  • ci gate
  • regression

rubric-grading-gate

A release gate for automatically-marked answer sheets. It grades against the rubric, then decides which of those grades are supported well enough to reach a student — naming the criterion, the reason, and the words on the page for every one it holds.

Why it matters. It checks the evidence for each answer, not the page as a whole. On 480 sheets carrying 156 wrong grades, a page-confidence threshold releases 378 with 128 wrong (66.14%); this gate releases 326 with 6 wrong (98.16%), and 97.4% of what a teacher opens is genuinely wrong. Six flags were written and two were cut by ablation — they add 49 sheets to the queue and catch nothing the other four miss. Standard library only, offline, seeded fixtures, 25 tests.

  • data quality
  • release gate
  • edtech
  • deterministic

llm-deprecation-radar

Static-scans a codebase for retired and soon-to-retire LLM model IDs, API surfaces, and parameters, then reports days remaining, blast radius, and the published replacement path. The deterministic demo catches 11 of 12 planted cases; the twelfth is a current-model negative control.

Why it matters. Every finding says how sure it is: 6 Confirmed, 2 Likely, 2 Unverified, and 1 runtime-data-required finding — plus registry freshness and a weekly CI deadline gate. Dynamic selection is never mislabeled as clean.

  • static analysis
  • llm ops
  • ci gate
  • deterministic

llm-cost-optimizer

A FinOps pass over an LLM usage log: reads a month of calls + a pricing table and reports how much spend is recoverable, with the dollars behind every recommendation. Model right-sizing + prompt-cache savings, runnable as a spend gate.

Why it matters. Turns "we should use a cheaper model" into a number: 38.9% of a synthetic month recoverable, the biggest win priced ($16.28 from moving classify off gpt-4o). Every figure derives from pricing.json — auditable, reproducible.

  • finops
  • llm cost
  • observability

ai-ready-repo-setup

Scans a repo into a CLAUDE.md / AGENTS.md an agent can follow — build/test commands, coding rules, prohibitions, CI smoke test, Definition of Done — then detects when it goes stale and scores readiness deterministically.

Why it matters. The value is not the generated file but drift detection + a deterministic readiness check that keeps the context honest, and says plainly what it does not measure.

  • ai-tooling
  • claude.md
  • verification

Data quality & load gates

12 projects

For data that loads without errors and is still wrong. Hire me to write the contract that stops a bad batch before it reaches a dashboard.

data-contract-demo Live

Load a CSV or JSON in the browser and get contract violations and drift signals against an editable, inferred contract. Nothing is uploaded — every check runs on the page.

Why it matters. The fixtures are synthetic controls with planted faults, and the test suite includes a mutation control: removing each planted fault lowers the caught count by exactly one, so the checker is proved to be detecting rather than guessing.

  • javascript
  • data contracts
  • drift
  • browser-only

dbt-fault-drill

Plants one realistic fault at a time into a copy of your dbt seed — a duplicate key, an amount sent in cents, last year’s file loaded again — runs dbt build, and reports which faults your tests stopped and which went through silently. On the example project, the basic suite (unique + not_null on the keys) stopped 3 of 17; a full contract, written by the same author as the faults, stopped 14 of 17.

Why it matters. Each fault is measured by what it did to a number someone reads: one 45,000,000 amount got through the basic suite and moved revenue by +44,999,955. The three faults the contract still missed — an event re-ingested under a new key, an amount off by 10x, a stale file — are listed with the test that would stop each. Reviewed independently twice; every finding fixed with a regression test.

  • dbt
  • duckdb
  • data quality
  • fault injection

warehouse-quality-gate

A dbt contract — 15 tests on the staging models — that stops a bad batch before the mart is rebuilt from it. Duplicate keys, broken relationships, an absurd amount, a mixed-currency batch, an out-of-window date, and a status with the wrong case.

Why it matters. The sabotaged batch loads with zero errors and reports revenue of $4,905,051; a clean batch from the same generator reports $395,751. One row carries $4.5M of that. The contract fails 12 of 15 tests and skips the mart build; the clean batch passes 15 of 15.

  • dbt
  • duckdb
  • data quality
  • data contract

data-contract-guard

A schema + range + enum contract with drift detection for tabular & sensor (rosbag) batches. Catches the batch that still parses but whose meaning changed — a firmware unit switch (°C→°F), a new category, a sensor dropout — that a naive "did it parse?" gate ships.

Why it matters. Names the cause, not the symptom: 0 issues on the naive check vs 36 violations + drift (joint_temp mean 45→112, z=14.9) on the same batch. Runs as a CI gate; the clean batch passes untouched.

  • data contract
  • drift
  • rosbag
  • data quality

sheet-to-db-contract

A load contract for the spreadsheet that quietly feeds a production system — header drift including reordering, spreadsheet error literals arriving as text, type drift, ambiguous dates, leading-zero and scientific-notation mangling, structural noise like repeated headers and totals rows, primary-key integrity, and declared enums and ranges.

Why it matters. The import ran and loaded every row without an error. The gate returns 31 findings (17 error, 14 warning) on 24 rows: #REF!, #VALUE! and #N/A landed as text, a declared column vanished, a long id arrived as 1.23457E+11, and 03/04/2026 is reported as ambiguous rather than guessed. It runs on a CSV export, so it needs no Google credentials at all.

  • google sheets
  • csv
  • data contract
  • pre-load gate

metrics-contract

A single YAML contract for six business metrics, and a checker that diffs every dashboard against it — filter sets, timezone, currency handling, distinct-ness, expressions, rate denominators, unowned metrics and dead ones.

Why it matters. It does not just say "definitions differ" — it runs both definitions over the same data and prices the gap: net revenue 697,691 vs 755,388 (+8.3%), buyers 870 vs 3,119 (+258.5%). 10 of 10 drifts surfaced, 0 on the compliant dashboard.

  • duckdb
  • analytics
  • metric governance
  • dataviz

dag-guard

Static review for Airflow DAGs — 12 checks read the source with ast, so there is no Airflow install, no import, and no side effects. Catchup blast radius, dependency cycles, non-idempotent writes, unordered co-writers, missing timeouts, retries, SLAs and failure callbacks.

Why it matters. The sabotaged DAG is valid Python that renders normally in the UI — and queues 90,816 backfill runs on deploy. Every finding carries evidence, production impact and the fix. 12 of 12 caught, 0 false alarms on the clean DAG.

  • airflow
  • ast
  • orchestration
  • ci

geo-data-quality-gate

A load contract for GeoJSON before it reaches PostGIS — CRS declaration, coordinate range, ring closure, winding order, self-intersection, zero-area and too-short rings, hole containment, duplicate geometry, dimension and geometry-type conformance, plus a longitude/latitude axis-swap heuristic.

Why it matters. The file parses as valid JSON and the load job starts. The gate blocks the same file with 17 findings across 13 checks on 13 features: longitude 190.0 and latitude 95.0 outside their ranges, a ring that never closes, a polygon that crosses itself, a hole outside its shell, and 5 points that only fall inside the declared bbox once swapped — flagged as a heuristic, not a proof. Zero runtime dependencies: the segment-intersection and shoelace math is stdlib.

  • geojson
  • postgis
  • spatial data
  • data quality

fhir-quality-gate

Twelve semantic checks for FHIR R4 bundles that are already structurally valid — code-system mislabelling, ICD-10 without its decimal, UCUM units that contradict the LOINC code, implausible values, unresolvable references, naive timestamps and value-set casing.

Why it matters. Both bundles share identical base data. The drifted one silently drops 3 patients from a 46-patient measure denominator and neither run raises an error. 12 of 12 defect types caught, 0 findings on the clean bundle. Standard library only, no PHI.

  • fhir
  • healthcare
  • loinc
  • data quality

pii-guard

A pre-release gate for “anonymized” exports — column-agnostic PII detection, mask-uniqueness and hash-invertibility checks, k-anonymity over quasi-identifiers, and DSAR erasure completeness across every related table.

Why it matters. The masking review finds 0 issues and the export ships. The guard blocks the same file: 652 PII instances in three unlabeled columns, 799 masks for 799 subjects, 567 user_ids reversed from the delivery’s own send log, 87.5% of rows unique at k=1, and 16 surviving references to an erased subject. The properly anonymized variant passes at min k = 13.

  • gdpr
  • pii
  • k-anonymity
  • dsar

soc2-evidence

An evidence completeness gate for SOC 2 / ISO 27001 control sets — refresh windows, population coverage against the authoritative roster, artifact-type matching, and contradiction against the access log. Evidence engineering: it checks whether the pack is complete, it does not issue an audit opinion.

Why it matters. The readiness tracker reads 12 of 12 evidenced and books the auditor. The gate returns 14 findings across 7 controls on the same pack: evidence 501 days old against a 365-day window, a review covering 34 of 53 identities, a backup control with 0 restore tests, and 8 live production grants held by 5 leavers.

  • soc 2
  • iso 27001
  • compliance evidence
  • access review

repair-cases

Two synthetic before/after repair cases with the fault planted on purpose: a pipeline that completes without an exception while mapping reordered columns by position and misreading a changed encoding, and a static page carrying a blocking script, four font stylesheets, and an injected layout shift.

Why it matters. It says plainly what these cases are not. They demonstrate diagnosis on deliberately broken local code — no client system, no production data, no commercial website — and each fix is stated as holding for the planted fault, not as a general remedy.

  • diagnosis
  • before/after
  • etl
  • performance

Migration, reconciliation & dedup

8 projects

For moves between systems and records that must add up. Hire me to prove nothing was lost, doubled or silently merged.

migration-verify-demo Live

Compare a source and a target export and get key reconciliation plus field-level evidence — the step where a migration quietly breaks. Both files stay in the browser.

Why it matters. It does not decide whether the migration should be accepted. It produces the evidence a person signs off on, and says so on the page.

  • javascript
  • migration
  • reconciliation
  • 16 tests

migration-verify

A post-migration verification gate for the losses a row count cannot see — primary-key reconciliation plus column-aggregate parity between the legacy table and the new warehouse. Row churn that nets to zero, a lost numeric scale on the money column, a timezone shift onto the wrong calendar day, silent truncation and NULL→'' coercion.

Why it matters. The cutover check every team ships reports 12,000 = 12,000, 0 issues and signs off. The same pair yields 96 findings: 42 orders missing, 37 duplicated, $5,955.90 (10.3 bps) unaccounted, and 1,416 orders on the wrong day. A correct migration passes at 0.

  • data migration
  • reconciliation
  • cutover
  • data quality

odoo-migration-gate

A pre-flight gate for a legacy export on its way into Odoo — required fields, dangling many2one references, duplicate and near-duplicate external IDs, selection values outside the allowed set, non-ISO dates, thousands separators in decimals, credit-note sign convention, and header totals that disagree with their own lines.

Why it matters. The import wizard reports the file is readable and the load begins. The gate returns 16 findings across 9 checks on the same export: a 99.00 gap between amount_total and the sum of its lines, an EUR invoice with no rate against a USD company, two references pointing at partners that are not in the export, and ‘15/01/2026’ parsed as a date. A clean export passes at 0.

  • odoo
  • erp migration
  • data contract
  • referential integrity

ecommerce-order-reconcile

A three-way reconciliation for an online store — orders against fulfillments against payouts. Paid but never shipped, shipped but never paid, quantity mismatches, payout shortfalls attributed to fee drift or an unrecorded refund, duplicate charges separated from legitimate repeat purchases, refunds with no restock, and currency or cent-level rounding breaks.

Why it matters. All three dashboards agree they are fine. The reconciliation returns 11 findings across all 8 checks and names $80.90 of payout shortfall — $0.90 at 62 bps explained by fee drift, $80.00 at 8,290 bps unexplained by any recorded refund. Money is Decimal end to end, with a test asserting no float ever touches it. A clean set with a real repeat purchase and a real partial fulfillment passes at 0.

  • shopify
  • stripe
  • reconciliation
  • decimal

statement-to-ledger

Bank and card statements into a reconciled ledger — and a refusal, in plain English, for the ones that do not add up. Importers are scored on rows parsed; this is scored on rows wrong: silently wrong rows 26 → 0, silently wrong money $62,832.40 → $0.00, unreconciled $8,905.20 → $0.00, corrupt sources caught 0/5 → 5/5.

Why it matters. It accepts fewer statements on purpose — 11 of 16 instead of 16 of 16 — and holds the rest rather than parsing them wrongly. Clean sources held: 0. The refusals are targeted, not blanket caution.

  • reconciliation
  • finance
  • refusal
  • decimal

dupe-merge-gate

Decide what your duplicates are before you delete any of them. The three-line DELETE … WHERE id NOT IN (SELECT MIN(id) …) against a gated path over eleven cases: facts lost 7 → 0, rows destroyed 6 → 0, rows fabricated 4 → 0, duplicates left behind 4 → 0, deletes with no undo 13 → 0, blocking operations in the plan 22 → 0.

Why it matters. It holds conflicting records instead of merging them. Rows that share an email but disagree on external_id are not one person, and the gate says so instead of picking a winner — 3 held for a human, 1 migration refused, and 0 false alarms on the sets that really were duplicates.

  • deduplication
  • migration safety
  • reversible
  • sql

crm-contact-hygiene

Deduplication and merge audit for contacts a CRM already holds — blocking, per-field similarity, a three-way verdict, declared conflict-resolution rules, and an audit log that names which record won every field and why. Nothing is looked up externally; the tool reads one file and writes cleaned output.

Why it matters. Blocking cuts 153 naive pair comparisons to 7 (95.4% saved). On 18 records it merges 3, leaves 15, and puts 2 pairs in a review queue rather than guessing — one because two rules disagree on a field, one because the score lands in the grey band. A father and son at one address, and two employees sharing a switchboard number, are not merged; that guard is a test.

  • crm
  • deduplication
  • record linkage
  • merge audit

plan-vs-actual-recon

A plan/spec with tolerances vs a measured inspection log → a reconciliation report: matched items, missing, extra, and out-of-tolerance deviations, each with the numbers.

Why it matters. Structured reconciliation, not CAD/vision; ambiguous matches are surfaced for a human instead of being silently auto-resolved.

  • qa
  • reconciliation
  • tolerance

Pipeline monitoring & scraping

8 projects

For scheduled jobs and scrapers that fail quietly. Hire me to add the checks that notice a missed run or a site that changed underneath you.

pipeline-heartbeat

An offline run-ledger monitor for scheduled pipelines — liveness, missed schedules, empty success, median+MAD volume and runtime bands, schema and watermark drift, partial sources, duplicate batches, persistent finding state, and a period digest.

Why it matters. A pipeline can exit 0 while emitting 0 rows, or stop running without throwing anything. The committed fixture catches 8 of 8 planted fault patterns while the healthy control produces 0 findings.

  • data pipeline
  • monitoring
  • median+MAD
  • data quality

run-ledger-hardening-kit

A run-tracking table that has never once said “failed” is not evidence that nothing failed. 282 runs across four pipelines, every row marked success or still running, reconstructed from what the runs left behind: silent failures found 4/10 → 10/10, false accusations 4 → 0, duplicate alerts 3 → 0, runs that never ran 0/2 → 2/2.

Why it matters. It has an undecided column. Three runs come back honestly undecided rather than guessed at — and of those three, zero were really broken. The naive baseline is not invented: it is the three sensible rules somebody writes the afternoon they are asked whether the pipelines are healthy.

  • observability
  • silent failure
  • alerting
  • evidence

operations-canary

Scheduled, offline evidence for a data operation that may drift after delivery. Replays four committed weeks of an order feed through six stages — collect, normalize, declared checks, change record, alert payloads written but not sent, and exports a client can open without this codebase.

Why it matters. It does not overclaim. The README states plainly that the workflow badge is not evidence until the scheduled run actually completes in public — the same honesty the checks themselves apply. Standard library only; no network call, no alert sent, no credential persisted.

  • scheduled checks
  • drift
  • evidence
  • github actions

automation-run-audit

An audit of a no-code automation’s own execution history — for the runs that reported success and still lost data. Silent attrition at a filter step, duplicate delivery from a retry without idempotency, schedule gaps, retry storms whose eventual success hides the failures, success rows carrying an error payload, and events applied out of their own order.

Why it matters. The platform dashboard shows the scenario active and the success rate in the high nineties. The audit returns 9 findings across all 7 checks on 30 runs: 90 items dropped by a filter at 3.00% against a 1.00% tolerance, one order delivered twice an hour apart, a 9-day window where the schedule never fired, and a ‘cancelled’ applied before its own ‘created’. It reads an exported run log — no account access, no API token.

  • n8n
  • make.com
  • zapier
  • automation reliability

scraper-canary

Seven invariants that fire when a site changes underneath a working scraper — selector hit-rate, per-field fill-rate, a median+MAD row-count band, type and enum stability, tag-skeleton drift, and freshness. Runs offline against committed snapshots of a fictional catalogue.

Why it matters. The expensive scraping failure is not a crash — it is a job that keeps succeeding while the prices come back empty. The report states how many planted breakages were caught, how many were missed, and how many were false alarms.

  • web scraping
  • drift
  • monitoring
  • data quality

site-watch-diff

Tell me what changed on this page — not what moved on it. Change watchers fail in both directions at once, and both are quiet: over the same 12 material changes, missed changes 7 → 0 and false alarms 40 → 0, with unusable snapshots caught 0/5 → 5/5.

Why it matters. The baseline is not a straw man — it is this same pipeline with every gate switched off, which is exactly what a text-based watcher does. And it catches the changes that never appear on screen: a link’s destination, a form’s post target, a robots directive.

  • change detection
  • monitoring
  • ablation
  • offline

ai-product-extractor

Messy product listings → clean, validated, de-duplicated data. LLM + heuristic extraction with a per-row confidence score and human-review flags.

Why it matters. Per-row confidence and review flags — you know exactly which rows to double-check.

  • etl
  • data extraction
  • ai-automation

hardware-parts-data-pipeline

Multi-source scraper + unified-schema normalizer for robotics/electronics parts (Pololu + Adafruit, 380 products), with idempotent ETL into SQLite.

Why it matters. Unified schema + idempotent loads — re-runs don't duplicate or corrupt the store.

  • web scraping
  • etl
  • sqlite

Documents, RAG & extraction

7 projects

For document Q&A and extraction that has to cite its source. Hire me to measure retrieval, refusals and which extracted rows can be trusted.

rag-copilot

A document copilot that cites what it says, refuses when the evidence is not there, and ships the test that proves both. Same corpus, same 33 questions, same scorer: baseline 12/33 → harness 32/33, including 4/4 on superseded editions, 4/4 on regional scope, and 4/4 on questions the corpus cannot answer at all.

Why it matters. The control row is the honest part — both systems answer all 8 un-trapped questions, so the gain is not bought with refusals. The baseline is what most demos ship: it invented an answer to every unanswerable question and returned two confident numbers from a policy edition retired on 2025-12-31. Runs in Docker with network_mode: none.

  • rag
  • citations
  • refusal
  • ablation

docs-chatbot-kit

A support bot’s mistakes do not live in its sentences — they live between them. The layer between a model’s draft reply and the person reading it, scored over 17 conversations / 32 turns: conversations never handed off 4 → 0, replies that must not go out but were sent anyway 4 → 0, good conversations interrupted 3 → 0.

Why it matters. Single-turn evaluation passes every failure in this repository. Read turn 6 of B-06 alone and it is a correct, well-cited sentence; read it after turn 1 and it is the bot telling the same person the opposite of what it just said. The naive baseline — a keyword escalation list — escalates the wrong half and misses every cross-turn failure.

  • support automation
  • escalation
  • multi-turn
  • guardrails

rag-grounded-qa

Grounded document Q&A that cites its sources, refuses when unsure (no hallucinations), and ships with an evaluation harness — recall@k, refusal accuracy, and grounding rate.

Why it matters. Not naïve RAG: the durable value is citation, refusal, and evaluation — scored, not hand-waved.

  • rag
  • retrieval
  • eval

rag-multidoc-crosscheck

Cross-document Q&A that surfaces every source's claim and flags contradictions — old vs new policy, US vs EU handbook — routing conflicts to human review instead of hiding them.

Why it matters. In policy & compliance, a hidden contradiction between sources is the expensive failure. Consensus/conflict detection is the trust layer.

  • rag
  • compliance
  • conflict-detection

rag-hybrid-rerank

Hybrid retrieval — BM25 + semantic, fused with Reciprocal Rank Fusion and reranked — with an evaluation harness that measures the lift (recall@k, MRR, nDCG) over any single retriever.

Why it matters. Retrieval quality is where RAG lives or dies. Here it is a measured number across four configs, not a claim.

  • rag
  • hybrid-search
  • rerank

invoice-to-structured

Documents (invoices, purchase orders, bank statements — PDF or scanned) → structured JSON/CSV through a deterministic verification engine: line math, balance reconciliation, totals, required fields, confidence scores, and a cross-document human-review gate. Optional AWS Textract path for scans.

Why it matters. The trust layer is the point — it flags which rows you can believe, not just extract-and-hope.

  • idp
  • textract
  • reconciliation
  • validation

screenshot-to-sheet

Screenshots → a spreadsheet, with a verification layer between them: per-field confidence gates, type and pattern checks, and a subtotal + tax = total reconciliation that catches OCR errors which parse cleanly. Ships with a Google Apps Script + Cloud Vision deployment path, so the client never touches the code.

Why it matters. OCR always misreads something. The number that matters is how many misreads were caught before the clean sheet — and the report states how many slipped through.

  • ocr
  • apps-script
  • google-sheets
  • validation

Workflow & intake automation

3 projects

For lead and request intake that routes itself. Hire me to build routing that sends uncertain cases to a person instead of guessing.

ai-intake-qualifier

Raw lead / form submissions → normalized, classified, and scored (hot / warm / cold) → a staff summary and reply draft → routed by a verification gate: trustworthy leads auto-flow to the CRM, uncertain ones go to a human-review queue. Runs fully offline (deterministic) or with OpenAI.

Why it matters. An LLM can label a lead; the value is knowing which to auto-route and which a human must see first — unreachable "hot" leads and ambiguous spam get caught, not dropped.

  • lead scoring
  • routing
  • qa gate

intake-agent-n8n

An n8n workflow that scores incoming inquiries with deterministic rules and routes them through a verification gate: complete leads land in a CRM-ready lane, uncertain ones stop in a human-review queue with explicit reasons, spam is rejected. Ships with the workflow JSON, a 24-case golden fixture set, ten unit tests and run screenshots — and the same scoring logic runs standalone with zero dependencies.

Why it matters. Anyone can wire n8n nodes together; the value is a gate that refuses to guess — every routing decision replays from committed fixtures (ALL PASS 24/24), and the executed branch was verified against the execution log, not just the response payload.

  • n8n
  • workflow automation
  • routing
  • qa gate

intake-router-multisource

Leads arrive from four differently-shaped sources (web form, Facebook Lead Ads, inbound email, CSV) → normalized to one schema → cross-source identity resolution merges the same person across channels → deterministic routing / SLA → a review gate for anything uncertain. Runs fully offline (deterministic) or with OpenAI.

Why it matters. Merging is a risk, not a convenience: only an exact email / phone match auto-merges — same name + same company but different contacts is flagged for a human, never silently combined. Full provenance and audit trail on every merge.

  • identity resolution
  • dedup
  • routing
  • qa gate

Analytics, forecasting & ML

13 projects

For reports and models people make decisions from. Hire me to build analysis that states its uncertainty and flags its own anomalies.

wnba-propboard Live

A dark, interactive WNBA player-prop board — sortable table with per-player sparklines, book line vs. projection with over/under coloring, and a click-through player-detail panel. Deployed as a standalone Vercel page that auto-updates on every push.

Why it matters. Data layer verifies every field against source (wnba.com / API) before it reaches the board; sample data shown in the demo is labeled as such.

  • html
  • dataviz
  • vercel
  • github ci/cd

analytics-pipeline

Messy CSV → cleaned data (with a quality log) → KPIs and month-over-month analysis → an automated report with charts. Includes anomaly detection: outliers, revenue mismatches, MoM spikes.

Why it matters. Data-quality checks and anomaly flags are baked into the report — not a pretty-but-blind dashboard.

  • pandas
  • reporting
  • anomaly

analytics-scheduled-report

A raw event/transaction log → scheduled daily/weekly reports with period-over-period deltas, threshold + statistical-outlier alerts, and idempotent re-runs.

Why it matters. Every alert carries its triggering rule and measured value — flags for a human to confirm, not a black-box "something's wrong".

  • analytics
  • scheduling
  • alerting

analytics-significance

Experiment/metric data → honest significance testing: two-proportion z-test and Welch's t-test (stdlib), confidence intervals, multiple-comparison correction (Bonferroni + Benjamini-Hochberg), and minimum detectable effect.

Why it matters. "Not significant" is reported with the effect it *could* have detected, not as "no effect"; family-wide correction flips naive false positives.

  • statistics
  • ab testing
  • analytics

demand-forecast-pipeline

A demand time series → forecasts with prediction intervals from a rolling-origin backtest, honest error metrics, and an explicit "when NOT to trust this" section.

Why it matters. Uncertainty is the product: intervals come from observed backtest residuals and widen with horizon — no accuracy overclaim, stdlib-only math.

  • forecasting
  • time series
  • ml

ml-churn-pipeline

Customer-churn prediction with an honest evaluation: leakage-safe splits, a baseline to beat, 5-fold cross-validation, and a model card that states the limits. Logistic regression implemented in NumPy.

Why it matters. Leakage guard + baseline + CV + model card = the difference between a demo notebook and something you can rely on.

  • numpy
  • ml
  • model card

data-mining-retail

Market-basket association rules (Apriori, lift-based) + RFM customer segmentation — surfacing real cross-sell signals and at-risk customers, with explainable segments.

Why it matters. Lift-based real signal (not just co-occurrence) and segments a marketer can actually act on.

  • apriori
  • rfm
  • segmentation

retail-sequence-mining

A transaction log → sequential purchase patterns, cohort retention, and churn associations, with Wilson intervals so small groups aren't over-read.

Why it matters. Associations are labeled correlational, carry counts + confidence intervals, and recommend an experiment before anyone acts on them.

  • data mining
  • retail
  • churn

odds-consensus

De-margined fair prices and a cross-book consensus that refuses to merge markets that are not the same. One schema for five books, the margin stripped the same way everywhere, and a hard rule for when two markets are actually the same market — before anyone claims a book is off the market by 9.85%.

Why it matters. It records what it declines: skipped.json and devig_skips.json record what was refused and why, so the consensus is a statement about markets that were genuinely comparable. No keys, no network, byte-identical fixtures.

  • normalization
  • de-vig
  • consensus
  • reproducible

wnba-player-vs-defense

Reproducible Python pipeline building a WNBA player-vs-defense info sheet — player game logs joined to opponent defense, every field verifiable against source.

Why it matters. The data engine behind the PROPBOARD demo — source-verifiable, reproducible.

  • sports data
  • pipeline

mlb-matchup-infosheet

MLB game logs → a reproducible batter-vs-pitcher info sheet: season/L5/L10, home-away and vs-hand splits, and head-to-head history, every number traceable to its sample size.

Why it matters. Information, not betting advice — no edge/EV claims and small samples flagged. Same reproducible pipeline as the WNBA sheet, second sport.

  • sports data
  • pipeline
  • mlb

field-inspection-report

Site photo → object detection → checklist verification → a PASS / REVIEW / FAIL report with annotated images. A confidence gate routes uncertain calls to human review; the detector is swappable (offline or a production vision model).

Why it matters. The detector is replaceable; the verification and human-review layer is what matters.

  • computer vision
  • numpy
  • qa gate

site-progress-report

Tagged site-photo metadata (EXIF) + an inspection checklist → a construction progress report: percent complete per area/phase, a timeline, and missing / out-of-sequence / stale-area detection.

Why it matters. Metadata + checklist only — never claims to see image content; answers "is every required shot present, in order, and recent", not "is the work good".

  • automation
  • construction
  • reporting

Robotics data

9 projects

For robot learning teams. Hire me to audit a dataset before you train on it or pay for it: bags, episodes, migrations and policy evaluation.

robotdata-pipeline

Raw operator/rosbag logs → a training-ready ML dataset (RLDS/LeRobot-style) with episode-level splits, a data contract that fails the build on bad data, and an RLDS-vs-LeRobot format recommendation.

Why it matters. The "make it" that follows an audit's "it's broken": a contract violation blocks a bad dataset from being produced instead of silently shipping one.

  • robotics
  • etl
  • data contract
  • robotops

ros2-bag-data-audit

A (synthetic) ROS2 rosbag export of robot telemetry → a structured anomaly audit: per-topic publish-rate deviations, dropout gaps, header-vs-receive clock skew, out-of-range sensor values, TF frame gaps, and a cmd_vel-vs-odom stall signal — organized by how sure the evidence is.

Why it matters. Findings are split into Confirmed evidence / Likely causes / Unverified hypotheses / Additional data required — it narrows candidate causes and says what data would confirm them, never asserting a root cause from a single bag.

  • ros2
  • rosbag
  • robotops
  • anomaly detection

lerobot-dataset-migrate

A LeRobot v2.1 dataset → a validated migration to the v3 layout: documented field/layout mapping, then integrity checks (frame counts, features, index continuity, stats) so nothing is silently dropped.

Why it matters. Migration + validation, not a black box — it models the publicly documented format and reports what it preserved vs what still needs checking against real datasets.

  • robotics
  • lerobot
  • data migration
  • robotops

dataset-curation-dedup

A robot/ML dataset manifest → near-duplicate detection (Hamming/cosine on provided hashes), quality gates, a coverage/imbalance report, and a curated manifest with a justified rejection log.

Why it matters. Every drop has a reason and borderline duplicates go to human review, never silent auto-deletion; it operates on manifests + hashes, not raw pixels.

  • robotics
  • dataset
  • dedup
  • robotops

crossembodiment-align

Multiple robot demonstration datasets (different embodiments, action spaces, control rates, gripper conventions) → a mixability assessment: normalize, flag frequency/gripper conflicts and missing embodiment metadata, and a per-pair mix / mix-subset / do-not-mix report.

Why it matters. The metadata needed to mix robot datasets is often missing — even NVIDIA's BridgeData v3 has an empty robot_type. This finds it from manifests, trains nothing, and every verdict cites its measured reasons.

  • robotics
  • cross-embodiment
  • dataset
  • robotops

policy-eval-harness

Robot policy rollout logs → success rate with Wilson confidence intervals, per-task breakdowns, and A/B comparison (difference CI + two-proportion test), with small-sample flags.

Why it matters. Same training loss ≠ same real success rate: it compares measured success with uncertainty, and never reports a rate without its n and CI.

  • robotics
  • evaluation
  • statistics
  • robotops

ros2-ci-bridge

A ROS2 colcon workspace → a reproducible Docker build + GitHub Actions CI (colcon build/test) → an offline Python layer that parses the build/test logs into an evidence-graded build-health report.

Why it matters. Findings are split into Confirmed evidence / Likely causes / Unverified hypotheses / Additional data required — the heavy build runs in CI, the demo verifies the analysis layer offline, and reproducibility is reported as signals, never a false "reproducible: yes".

  • ros2
  • docker
  • ci/cd
  • robotops

ros2-urdf-smoke-test

A robot URDF → static structural audit (link/joint tree integrity, missing meshes, inertial sanity, joint limits) → a four-section smoke-test report.

Why it matters. Findings are split into Confirmed / Likely / Unverified / Additional data — a pass means "likely loads", never "works" or "safe".

  • ros2
  • urdf
  • robotics
  • robotops

domain-rand-sweep

A parameter-space spec → a domain-randomization sweep (grid / random / Latin-hypercube) plus coverage and pairwise-gap analysis, ready to feed a simulator.

Why it matters. Designs and analyzes the sweep only — it runs no simulator and makes no sim-to-real transfer claim: broad coverage is not evidence of real-world robustness.

  • robotics
  • domain randomization
  • sim
  • robotops

All 79 projects are public on github.com/jigonyoo and MIT-licensed. 73 of the 79 contain Python test files, and 65 pass at least one test in a throwaway environment where only pytest is installed and network access is blocked (python -m pytest --continue-on-collection-errors), measured 2026-09-26 for the first 76, 2026-09-30 for bad-batch-gate, and 2026-10-04 for dbt-fault-drill and replay-twice. Sample projects use synthetic or sample data where noted; what they demonstrate is the pipeline and its checks, not a specific dataset.

Trust

How I handle your code and data

You are giving a stranger access to your system. These are the rules I work by on every project.

  • NDA on requestI sign your NDA before you share anything.
  • Read-only accessI ask for the least access the job needs, usually read-only. No write access to production.
  • Masked samples are fineMasked, sampled or synthetic data works. I tell you up front what the check needs.
  • Deleted when we finishYour code and data are deleted at the end of the project, and I confirm it in writing.
  • Your findings stay privateNothing from your project appears on this site or anywhere else without your written permission.
  • Replies within 24 hoursOn weekdays, Korea time (UTC+9). Everything stays in writing, so nothing gets lost.

About

About Jigon Yoo

I'm Jigon Yoo, a freelance engineer in South Korea. I work on the part of AI and data systems that is easy to skip: checking whether the output is right, and saying plainly when it is not.

Most of what I build is a test bench. I plant the fault, measure whether a system catches it, and publish the number, including the cases where my own checks were wrong.

Focus
LLM agents and tool calls, evaluation and graders, data pipelines and load contracts
Stack
Python, SQL, dbt, DuckDB, SQLite, MCP, n8n, GitHub Actions
Location
South Korea (UTC+9), working remotely with teams anywhere
Style
Async and in writing: scope, questions, updates and the final report
Track record
100% Job Success · 10 jobs on Upwork

Contact

Tell me what's going wrong. I'll tell you how I'd measure it.

Email is the fastest way to reach me. I reply within 24 hours on weekdays (Korea time), with questions or a fixed-price quote.

jigondaniel1224@gmail.com Open email app

A useful first email covers

  1. What the system does, in a sentence or two
  2. What went wrong, or what you are about to change
  3. What you can share: repo, sample data, logs
  4. Any deadline

Prefer to hire through a platform? I'm on Upwork and Contra.