Services
I am a research engineer and technical lead with 15+ years across storage firmware, distributed systems, and mainnet operations. Engagements are hands-on: I read code, look at telemetry, and work with your engineers. You get written findings you can act on whether or not we keep working together.
AI-assisted engineering
LLM coding tools look great on demos and are unreliable by default in real codebases, because a model is only as good as the context and the verification around it. I work on making them productive in an engineering organization: context engineering, repository guidance files, AI review systems, agent workflows with verification built in, and measurement of actual throughput rather than the feeling of productivity. At Caffeine.ai I built the custom LLM PR-review system that became the team’s main review system, and an AI-native ticket-to-PR loop — plan, implement, self-review, fix, update ticket, respond to reviewer comments — that sustained 5–15 meaningful PRs per day while burning down tech debt.
What you get: repository guidance and context engineering fitted to your codebase, review automation that enforces your standards before a human looks, agent workflows for the work worth automating, and numbers that show whether any of it is paying off.
Production reliability
Production systems fail in ways that are obvious in hindsight and missing from the dashboards — until someone fixes the dashboards. I work on observability, incident reduction, rollout safety, migrations without downtime, and on-call and operations maturity. At DFINITY I led reliability for the Internet Computer, a 1,400+ node HA network: critical incidents went from double digits per year to 1–2 while usage grew by orders of magnitude and the network shipped hundreds of upgrades per year. I built observability across a multi-AZ Kubernetes estate, and with two colleagues rewrote the CI system, cutting merge times from 1–2 days to 30–60 minutes.
What you get: working telemetry with alerts people trust, a measured drop in incident count and severity, rollouts you can reverse, and migrations that do not need a maintenance window.
Storage and distributed systems
Storage is where systems are quietly won and lost: latency budgets, durability, failure modes, and the trade-offs that only show up under real load. I do architecture reviews and design work for storage engines, data platforms, and the distributed systems underneath them, from flash firmware behavior up to cluster-level design. At IBM Research I was one of the core FTL firmware developers for the IBM FlashCore Module. Later I architected an analytics platform ingesting real-time telemetry from 11,000+ enterprise storage systems. That grounding comes from a research career: 16+ peer-reviewed papers and 60+ patents.
What you get: an architecture review that covers the whole stack, written findings with concrete recommendations, and a design your engineers can actually implement.
Working together
Engagements range from a short assessment with written findings to ongoing advisory. If any of this sounds like your problem, find me on LinkedIn.