Vikrant Singh

IT Operations and Engineering

I've led teams, and I build hands-on: useful systems for businesses, including my own.

For fifteen years I've been accountable for IT services reaching people and staying up.

At Citi, through Virtusa, I ran delivery across 22 Trade and Transaction Services programs in North America, from design to go-live, including the infrastructure for all 22. Root-cause work on two trade applications brought repeat incidents down.

For the Government of Ontario, through CompuCom (2018–2021), I led a 15-person team supporting a 24x7 estate of 3,000+ servers, and was the subject-matter expert on its ServiceNow–Remedy integration. I was the person the client called on Sev-1 and Sev-2, and I ran the response through restore. We moved a large share of that volume onto automated resolution and shortened restore time by tightening the integration path and the runbooks around it.

What I hold a team to. Work isn't done without tests, monitoring, a way to roll back, and a named person who gets paged. Status means remaining work and risks, not percent complete. When something breaks at 3 a.m., I pick up. When something breaks, the people affected hear the truth first, and the fix changes what happens next time.

What I'm building toward. Owning a product team end to end: the build as well as the run. How I run this site as a product →

A short path

If you only have a minute:

  1. Live reliability: this site's own SLOs and error budgets (still early readings) and recent failures.
  2. When the write budget ran out: a load test of mine used up the day's database writes, and what changed after.
  3. Incident desk: alerts land as public issues; an agent proposes, and a human authorizes every decision after that.
  4. Game day 1: a planned fault, already run, through detect, triage, and restore.

Proof you can open

Writing

Notes from running this site, and short essays on putting AI into operations without losing track of who decides. Read the notes and essays →