Ashik Saeed
All case studies

Production case study

CloudlyMELT — AI Infrastructure Observability

Building CloudlyMELT with the CloudlyIO team: an AI-native observability platform that connects network-fabric, GPU/system, and AI-workload telemetry to explain infrastructure failures across the stack.

AI InfrastructureObservabilityIncident IntelligenceOpenTelemetryGoKubernetesApplied AI
CloudlyMELT — AI Infrastructure Observability project preview
Client
CloudlyIO
Engagement
Jan 2026 – Present
My role
Lead engineer driving RCA, incident intelligence, validation, and delivery hardening

3 layers

network fabric, GPU/system, and AI workload

5 stages

from symptom detection to root-cause evidence

2 use cases

training throughput and inference latency RCA

01 · The brief

The assignment and its real constraints

When a distributed AI workload slows down, the visible symptom may be far from the cause. Operators need to connect fabric congestion, GPU behavior, storage, orchestration, and model-serving signals without manually stitching together several observability products.

01

Correlate signals across infrastructure layers without presenting synthetic or incomplete evidence as a production finding.

02

Complement existing Prometheus, Grafana, OpenTelemetry, and incident-response tooling rather than requiring a full replacement.

03

Support Kubernetes, cloud, on-premise, tenant isolation, and degraded dependencies while keeping incident contracts stable.

04

Make controlled demonstrations reproducible while maintaining an explicit evidence boundary for real hardware and runtime validation.

02 · My mandate

What I owned

  • Drove major parts of the inference-latency and training-throughput RCA delivery, from scenario design and signal normalization through dashboards and validation gates.
  • Designed and implemented the durable Incident Intelligence contract, persistence model, tenant boundaries, lifecycle behavior, and API tests.
  • Hardened deployment and operational paths across Helm, Docker, Terraform, CI, documentation, and live-readiness tooling.
  • Work with the product and engineering team to convert an evolving observability thesis into reviewable, customer-facing increments.

03 · Technical judgment

Decisions that shaped the system

01

Correlate the layers instead of adding another dashboard

The product models relationships between network fabric, GPU/system telemetry, and AI workloads so operators can follow a causal chain rather than compare disconnected charts.

02

Preserve evidence with every incident

Stable incident IDs, immutable initial evidence, refreshable current evidence, and lifecycle persistence make RCA findings reviewable beyond one detection cycle.

03

Treat unavailable context as a first-class state

Investigation sources expose provenance, freshness, and explicit degraded or unavailable states instead of filling gaps with invented context.

04

Meet operators inside their existing stack

OpenTelemetry, Prometheus, Grafana, Kubernetes, and versioned APIs keep CloudlyMELT complementary to established observability and incident workflows.

04 · Execution

From discovery through production

  1. Step 1

    Translate priority operational failures into observable symptoms, candidate root causes, required signals, and acceptance evidence.

  2. Step 2

    Normalize telemetry from AI runtimes and infrastructure receivers into a shared cross-layer correlation model.

  3. Step 3

    Exercise the RCA pipeline through controlled scenarios, automated checks, dashboards, and real-telemetry readiness gates.

  4. Step 4

    Persist the resulting incident, evidence, lifecycle, and tenant context behind a versioned external contract.

  5. Step 5

    Iterate with the team on investigation context, operator experience, deployment hardening, and customer-validation paths.

System delivered

The production surface

Cross-layer correlation across network fabric, GPU/system, and AI workloads
Training-throughput and inference-latency root-cause analysis
Durable tenant-scoped incident lifecycle and evidence APIs
vLLM and AI-runtime telemetry normalization and validation
Kubernetes, cloud, on-premise, and air-gapped deployment paths
Integration with established observability and incident-response stacks

05 · Result

Production outcomes

Delivered multi-stage correlation and controlled RCA paths spanning ten infrastructure and AI-workload failure scenarios
Productionized the inference signal path and built repeatable EC2, vLLM, Prometheus, and Grafana evidence gates
Built durable, tenant-scoped Incident Intelligence APIs with stable incident IDs, lifecycle operations, evidence retention, and restart persistence
Currently extending incidents with bounded dependency, blast-radius, ownership, recent-change, and provenance context

Evidence & disclosure

CloudlyIO describes CloudlyMELT as an AI-native unified observability platform for GPU-scale infrastructure, spanning cross-layer correlation, root-cause analysis, Incident Intelligence APIs, OpenTelemetry pipelines, a Python training SDK, and Kubernetes deployment support.

Public scale and product claims are attributed to the linked source. Delivery details reflect my direct role; confidential client data and implementation details are intentionally omitted.

Need someone who can clarify the real constraint, make the technical trade-offs explicit, and stay accountable through production delivery?

Discuss a project