Work
Systems we've shipped
Six engagements, six different failure modes — a ledger that couldn't reconcile itself, a monolith that took down checkout on every deploy, a decade of inconsistent permission checks. Filter by discipline, or scroll through all of it below.
In detail
Each engagement, with the specifics.
Veridian Pay
A payments startup whose MVP ledger couldn't survive its own growth — rebuilt as a reconciliation-safe, audit-ready transaction system.
Challenge
Veridian's first ledger was built to get a demo in front of investors, not to survive real transaction volume. Balances were computed by summing rows at read time, there was no idempotency on retried payments, and a single slow query could lock the whole write path. It worked at 200 transactions a day. It did not work at 20,000.
Approach
We rebuilt the ledger as an append-only, double-entry system with idempotency keys on every write, and pre-computed balances maintained by an event stream rather than recalculated on read. Reconciliation against the payment processor became a scheduled job instead of a support team's manual spreadsheet. None of it was novel architecture — it was the boring, well-understood pattern the first version had skipped under deadline pressure.
Outcome
Veridian passed its Series A technical diligence without a single ledger-related finding, and the on-call rotation stopped getting paged for balance discrepancies entirely. The rebuild shipped without a single hour of payment downtime.
Harborline
A two-sided marketplace moved off a single overloaded VM and onto infrastructure that could survive a deploy without taking checkout down with it.
Challenge
Harborline ran its entire marketplace — API, background workers, search indexing, and a cron-driven payout system — on one large EC2 instance. Every deploy required a maintenance window, and a memory leak in the search indexer had, twice, taken down checkout along with it, because everything shared the same process supervisor.
Approach
We split the deployment surface — not the codebase, that rewrite wasn't the actual bottleneck — into independently deployable services running on Kubernetes, with resource limits isolating the search indexer from the request-serving API. Deploys moved from a nightly maintenance window to rolling updates with automated rollback on failed health checks.
Outcome
Harborline now ships to production multiple times a day without a maintenance window, and the search indexer has crashed twice since launch — without taking anything else down with it.
Northfield Health
A patient-scheduling platform's decade-old PHP core got a security and access-control overhaul without a disruptive full rewrite.
Challenge
Northfield's scheduling system had grown over nine years into a single PHP codebase with role permissions checked inconsistently across roughly 40 different entry points. A security review ahead of a hospital-system contract flagged this directly: there was no way to confidently state who could see what, because the answer was different depending on which file you read.
Approach
Rather than a ground-up rewrite — which would have meant a year with no new features — we built a centralized authorization layer that every entry point was migrated through one route at a time, verified against a test suite written specifically to catch permission regressions before the migration touched that route. Six months of incremental, measurable progress rather than one high-risk rewrite.
Outcome
Northfield passed its hospital-system security review on the first attempt, with the audit team specifically noting the consistency of the access-control model. The migration shipped without a single week of feature freeze.
Cargopath
A logistics startup's shipment-tracking data went from a 15-minute polling delay to sub-second updates with an event-driven pipeline.
Challenge
Cargopath's dispatchers needed to know where a shipment was right now, but the system polled each carrier's API on a schedule and wrote results into a table the dashboard re-queried every page load. At dispatcher-relevant scale, that meant status updates arriving up to 15 minutes late — long enough to miss a delay that actually mattered.
Approach
We replaced the polling loop with an event-driven pipeline: carrier webhooks, where available, and a tightened polling schedule fed a message queue, which drove both the database write and a live dashboard update over a websocket connection — so dispatchers saw a status change the moment it existed, not the next time a query happened to run.
Outcome
Shipment status now reaches the dispatcher dashboard in under two seconds instead of up to fifteen minutes, and Cargopath used the same event pipeline to add a customer-facing live tracking page without any additional backend work.
Classly
An edtech platform's dashboard load time dropped from 8 seconds to under 900ms without touching the product's feature set.
Challenge
Classly's teacher dashboard issued 40+ sequential database queries on every page load, several of them full-table scans that had been fast at launch and had quietly gotten slower as the student roster grew. Teachers — the platform's least patient users, logging in between classes — were the ones feeling it most.
Approach
We profiled the full request path rather than guessing, found the handful of queries responsible for the vast majority of load time, added the missing indexes and a Redis caching layer for data that changed infrequently, and batched the remaining sequential queries into two. No framework changes, no rewrite — just measuring before optimizing.
Outcome
Dashboard load time dropped from roughly 8 seconds to under 900 milliseconds, and Classly's support team reported a noticeable drop in "is the site down" tickets within the first week.
Opsdeck
A B2B SaaS team went from engineer-triggered manual deploys and no real on-call process to automated CI/CD and a rotation that actually worked.
Challenge
Opsdeck's deploys happened when an engineer remembered to run a shell script, usually late on a Thursday, and "on-call" meant whoever happened to see a Slack message first. There was no staging environment that reliably matched production, so most incidents were discovered by customers before they were discovered internally.
Approach
We built a CI/CD pipeline with automated tests gating every deploy, a staging environment provisioned from the same Terraform config as production, and a structured on-call rotation with defined escalation paths and a runbook written for the incidents that had actually happened before — not hypothetical ones.
Outcome
Deploys went from a manual, Thursday-night event to an automated process that happens whenever code merges, and mean time to detection for incidents dropped from customer-reported to alert-triggered within the first month of the new rotation.
Start a project
Got a system that worked at launch and is starting to crack?
Tell us what's breaking and where. If it's a good fit, you'll hear back from an engineer directly — not a sales rep.