IT Operations and Engineering
I've led teams, and I build hands-on: useful systems for businesses, including my own.
For fifteen years I've been accountable for IT services reaching people and staying up.
At Citi, through Virtusa, I ran delivery across 22 Trade and Transaction Services programs in North America, from design to go-live, including the infrastructure for all 22. Root-cause work on two trade applications brought repeat incidents down.
For the Government of Ontario, through CompuCom (2018–2021), I led a 15-person team supporting a 24x7 estate of 3,000+ servers, and was the subject-matter expert on its ServiceNow–Remedy integration. I was the person the client called on Sev-1 and Sev-2, and I ran the response through restore. We moved a large share of that volume onto automated resolution and shortened restore time by tightening the integration path and the runbooks around it.
What I hold a team to. Work isn't done without tests, monitoring, a way to roll back, and a named person who gets paged. Status means remaining work and risks, not percent complete. When something breaks at 3 a.m., I pick up. When something breaks, the people affected hear the truth first, and the fix changes what happens next time.
What I'm building toward. Owning a product team end to end: the build as well as the run. How I run this site as a product →
A short path
If you only have a minute:
- Live reliability: this site's own SLOs and error budgets (still early readings) and recent failures.
- When the write budget ran out: a load test of mine used up the day's database writes, and what changed after.
- Incident desk: alerts land as public issues; an agent proposes, and a human authorizes every decision after that.
- Game day 1: a planned fault, already run, through detect, triage, and restore.
Proof you can open
This site's reliabilityLive
Two service level objectives with live error budgets; an outside probe checks the site every five minutes. Two small games feed it real traffic: a 30-second runner measures what the browser experiences, and a multiplayer Ludo game measures how long every move takes on the server. I break it on purpose and publish what I find, and when a load test of my own took part of it down, I wrote that up too.
Live reliabilityPostmortem: game dayPostmortem: real incidentPlay the games
Incident deskLive
Real alerts from this site go to an AI that proposes triage and drafts updates. Drafts are never sent; resolving and closing are always a person's call, and since 1 Oct 2026 so is priority, in public GitHub issues. If the model is down, the incident record still works.
Scope: a Python engine running one site's incidents, not ServiceNow.
Watch a live game day (~9 min)Watch the walkthrough (~14 min)See it liveHow it works
Balance-BooksLive
A live bookkeeping site with a CRA readiness check: six questions, a score out of 12, and the areas to tighten first. If the visitor asks to talk, the contact form arrives with their score and gaps attached, so the context isn't lost. Bot protection with Cloudflare Turnstile, a serverless contact API, and business email through Cloudflare Email Routing.
Scope: a small business site; the check is a self-assessment, not tax advice.
Atlas Flow
A live TypeScript/Node write path: web form, server-side validation, then a row written to SharePoint in production.
Scope: a marketing site, not a system with millions of members.
Writing
Notes from running this site, and short essays on putting AI into operations without losing track of who decides. Read the notes and essays →
Résumé (PDF) · Email · LinkedIn · GitHub