Engineering assistant
Explore an agentOrchestrator
Receives a task, dispatches specialists and brings their findings together.
IoT Infrastructure Architect & Lead Engineer
I design cloud platforms for Culligan’s connected devices, and build the AI tools that help our engineers review, test and ship.
Fabrizio De Cicco · Turin, Italy · Open to remote
Receives a task, dispatches specialists and brings their findings together.
Architecture, reliability and engineering practice for Culligan's IoT platform.
10+years in infrastructure
~300,000fleet scale target
I designed and built a multi-agent engineering assistant for review, testing, release preparation and reliability. It is now used across the lead engineering team.
Explore the architecture
control
↓ dispatches to specialists
specialist agents
↑ all built on a shared foundation
shared
↕ connected to the toolchain via MCP
integrations · MCP
Each agent owns one workflow and carries only the context it needs — organised by workflow, not by role.
Across a growing set of repositories, code review, QA, release coordination, and reliability monitoring were largely manual — slow to perform, hard to keep consistent, and a constant pull on senior engineers' time.
I designed and built a set of specialist AI agents in Claude Code, each owning a single workflow and sharing a common library of reusable skills. The orchestrator composes them into dynamic workflows that fan work out in parallel and adversarially verify findings before anything is surfaced. They integrate with the team's existing toolchain through MCP, sit behind safety guardrails that gate anything destructive, and the read-only ones run as scheduled loops that catch drift and post a summary — never acting unattended. The newest layer points the other way — hooks and skills that sharpen the engineer's side of the loop: reframing vague prompts at session start, surfacing blindspots before unfamiliar work, and quizzing the author on their own change before it ships.
Rolled out to the lead engineering team and in daily use — seven agents sharing a 50-skill library, behind guardrails that gate anything destructive. Review quality has visibly improved and developer feedback has been consistently positive; the next iteration adds the numbers I want to manage it by: cycle time and defect leakage.
Modeled as a two-phase workflow: fan out → adversarially verify.
Equal task durations, separate capacity per stage and no coordination overhead. An illustrative model, not measured delivery time.
The decisions I come back to when building platforms and the teams behind them.
For connected products, the platform's job is to keep a fleet secure, updatable, and visible — not to ship features on day one. Device identity, over-the-air update, and end-to-end telemetry are the load-bearing walls; bolt them on later and you're rebuilding the foundation under a live fleet.
The cheapest time to standardize IaC modules, secret management, pipelines, and an SLO convention is at two repos, not twenty. Every week you wait, divergence compounds and migration turns political. Standardization isn't bureaucracy — it's the paved road that lets people stop reinventing plumbing.
AKS is right when you have many services, real scaling needs, and a team that can own the operational surface. For a handful of services it's a tax — managed container platforms ship the same outcome with a fraction of the ops. Complexity should be earned by load, not adopted for the résumé.
Chasing more nines than users actually feel just burns money and engineering time. Define the SLO from real user experience, then spend the error budget: ship faster while it's healthy, slow down and harden when it's not. Reliability and cost aren't opposites — the error budget is the dial that trades them on purpose.
AI agents are a force-multiplier on review, QA, and release prep — but anything destructive stays behind a human gate and a guardrail. Compose them into workflows that fan out and self-verify, give them a shared skills layer so conventions don't drift, and run the read-only ones on a schedule so problems surface on their own. The win is consistency and removed toil, not autonomy for its own sake.
A lead who hoards context becomes the bottleneck: the team stalls whenever they're unavailable, and burnout follows. My job is to make myself progressively unnecessary — mentor, document, pave the road — so the team moves faster than any one person could.
Templates I actually use — de-identified and free to take — plus the number behind every "we need more nines" conversation.
Each is a real .md in the /lab folder of this site's repo.
Pick an availability target and see how much downtime it actually buys you.
Azure · AKS · Container Apps · Kubernetes · Docker · Helm
IoT Hub · Event Hub · MQTT · CoAP · LWM2M · device identity · OTA
Terraform · Terragrunt · Bicep · Ansible · Azure DevOps · YAML pipelines
Datadog · Grafana · Prometheus · SLOs · error budgets · incident response
Claude Code · MCP · multi-agent orchestration · model tiering · TypeScript
CKA Certified Kubernetes Administrator (2022) · LFCS Linux Foundation Certified Sysadmin (2023) · AZ-104 Azure Administrator (2023) · AZ-400 DevOps Solutions (2023)
Architecture, platform reliability, or AI-assisted delivery — if any of that is on your plate, I'm happy to talk.