Production case study
CloudlyMELT — AI Infrastructure Observability
Building CloudlyMELT with the CloudlyIO team: an AI-native observability platform that connects network-fabric, GPU/system, and AI-workload telemetry to explain infrastructure failures across the stack.
- Client
- CloudlyIO
- Engagement
- Jan 2026 – Present
- My role
- Lead engineer driving RCA, incident intelligence, validation, and delivery hardening
3 layers
network fabric, GPU/system, and AI workload
5 stages
from symptom detection to root-cause evidence
2 use cases
training throughput and inference latency RCA
01 · The brief
The assignment and its real constraints
When a distributed AI workload slows down, the visible symptom may be far from the cause. Operators need to connect fabric congestion, GPU behavior, storage, orchestration, and model-serving signals without manually stitching together several observability products.
Correlate signals across infrastructure layers without presenting synthetic or incomplete evidence as a production finding.
Complement existing Prometheus, Grafana, OpenTelemetry, and incident-response tooling rather than requiring a full replacement.
Support Kubernetes, cloud, on-premise, tenant isolation, and degraded dependencies while keeping incident contracts stable.
Make controlled demonstrations reproducible while maintaining an explicit evidence boundary for real hardware and runtime validation.
02 · My mandate
What I owned
- ✓Drove major parts of the inference-latency and training-throughput RCA delivery, from scenario design and signal normalization through dashboards and validation gates.
- ✓Designed and implemented the durable Incident Intelligence contract, persistence model, tenant boundaries, lifecycle behavior, and API tests.
- ✓Hardened deployment and operational paths across Helm, Docker, Terraform, CI, documentation, and live-readiness tooling.
- ✓Work with the product and engineering team to convert an evolving observability thesis into reviewable, customer-facing increments.
03 · Technical judgment
Decisions that shaped the system
Correlate the layers instead of adding another dashboard
The product models relationships between network fabric, GPU/system telemetry, and AI workloads so operators can follow a causal chain rather than compare disconnected charts.
Preserve evidence with every incident
Stable incident IDs, immutable initial evidence, refreshable current evidence, and lifecycle persistence make RCA findings reviewable beyond one detection cycle.
Treat unavailable context as a first-class state
Investigation sources expose provenance, freshness, and explicit degraded or unavailable states instead of filling gaps with invented context.
Meet operators inside their existing stack
OpenTelemetry, Prometheus, Grafana, Kubernetes, and versioned APIs keep CloudlyMELT complementary to established observability and incident workflows.
04 · Execution
From discovery through production
Step 1
Translate priority operational failures into observable symptoms, candidate root causes, required signals, and acceptance evidence.
Step 2
Normalize telemetry from AI runtimes and infrastructure receivers into a shared cross-layer correlation model.
Step 3
Exercise the RCA pipeline through controlled scenarios, automated checks, dashboards, and real-telemetry readiness gates.
Step 4
Persist the resulting incident, evidence, lifecycle, and tenant context behind a versioned external contract.
Step 5
Iterate with the team on investigation context, operator experience, deployment hardening, and customer-validation paths.
System delivered
The production surface
05 · Result
Production outcomes
Evidence & disclosure
CloudlyIO describes CloudlyMELT as an AI-native unified observability platform for GPU-scale infrastructure, spanning cross-layer correlation, root-cause analysis, Incident Intelligence APIs, OpenTelemetry pipelines, a Python training SDK, and Kubernetes deployment support.
Public scale and product claims are attributed to the linked source. Delivery details reflect my direct role; confidential client data and implementation details are intentionally omitted.
Need someone who can clarify the real constraint, make the technical trade-offs explicit, and stay accountable through production delivery?
Discuss a project