Enterprise4 min read

Connect Your Production, Simulate Against It: The Auto-Calibration Engine

Import real metrics from Datadog, real incidents from Jira, real infrastructure from Terraform — and the simulation automatically matches your production behavior.

The Gap Between Simulation and Reality

Most architecture simulations use generic numbers. "Assume 100ms latency." "Assume 99.9% availability." "Assume 1000 req/s."

But your production doesn't work in round numbers. Your real P99 is 347ms. Your real peak is at 2:15 PM on Tuesdays. Your payment service fails 0.3% of the time, but only on international transactions.

What if your simulation used your real production data?

The Auto-Calibration Pipeline

PraxiRun connects to your existing monitoring, incident management, and infrastructure tools — then automatically calibrates the simulation to match your production behavior.

Step 1: Connect Your Monitoring

Connect your Datadog, New Relic, CloudWatch, or Prometheus instance. PraxiRun imports the last 30 days of:

  • Traffic patterns — when are your peaks? What's the diurnal curve?
  • Latency distributions — P50, P95, P99 per service, not averages
  • Error rates — which services fail? How often? With what patterns?
  • Resource utilization — CPU, memory, connection pools per service

Step 2: Connect Your Incident History

Connect Jira or ServiceNow. PraxiRun imports your incident records and extracts:

  • Failure frequency — "payment service had 3 incidents last month"
  • Mean Time To Recovery — "average MTTR is 45 minutes"
  • Root causes — "database connection exhaustion caused 60% of incidents"
  • Impact patterns — "incidents during peak hours affect 10x more users"

These become realistic failure scenarios in your simulation. Not generic "service down" scenarios — YOUR actual failure modes.

Step 3: Connect Your Infrastructure

Connect your Terraform state or Kubernetes API. PraxiRun imports:

  • Actual resource configuration — instance types, replica counts, HPA settings
  • Network topology — VPCs, subnets, security groups
  • Scaling rules — min/max replicas, CPU thresholds, cooldown periods

Your simulation now has the exact same infrastructure as production.

Step 4: Auto-Calibration

The calibration engine combines all three sources and produces a SimCalibrationProfile:

Traffic: diurnal pattern, base=45 rps, peak=380 rps (Tuesdays 2:15 PM)
Latency: P50=23ms, P95=89ms, P99=347ms (from Datadog)
Error rate: 0.3% (payment service, international only)
Replicas: 3 (from Terraform state)
Scenarios: "Payment timeout" (3/month, MTTR 45min)
           "DB connection exhaustion" (1/month, MTTR 90min)
Cost: $12,400/mo (from billing)
Availability: 99.92% (calculated from real error rate)

Your simulation now behaves like your production. Every latency, every error rate, every scaling behavior — calibrated from real data.

What You Can Do With a Calibrated Simulation

"What if we add a cache?"

Before: you guessed it would help. Maybe 30% improvement?

After: simulation shows P99 drops from 347ms to 52ms. DB utilization drops from 78% to 22%. Cache hit rate stabilizes at 73% after warmup. Cost increases by $45/mo.

You have the exact numbers. Not estimates — data.

"What if we lose an availability zone?"

The simulation kills all services in AZ-1a (matching your real Terraform topology). Traffic redistributes to AZ-1b and AZ-1c. Simulation shows:

  • 23 seconds of elevated error rate during failover
  • P99 spikes to 890ms for 45 seconds
  • Auto-scaling adds 2 replicas in 90 seconds
  • Full recovery in 2 minutes

You know exactly what happens. Before it happens.

"What if Black Friday traffic hits?"

Multiply your real peak traffic by 5x. Simulation shows:

  • Database connection pool exhausts at 3.2x current peak
  • Payment service becomes the bottleneck at 4.1x
  • Cache hit rate drops to 45% (cold cache for new product pages)
  • Breaking point: 4.8x current peak

You know your limits. And you know exactly which component to scale.

30 Connectors

PraxiRun connects to your entire stack:

| Category | Connectors | |---|---| | Monitoring | Datadog, New Relic, Dynatrace, CloudWatch, Prometheus, Elasticsearch, Grafana | | Incidents | Jira, ServiceNow, PagerDuty, OpsGenie | | Infrastructure | AWS, Azure, GCP, Kubernetes, Terraform, Istio | | Physical | MQTT, ROS 2, OPC-UA, MAVLink, HL7 FHIR, BACnet | | CI/CD | GitHub, Jenkins | | Communication | Slack, Email (Resend) |

Connect once. Calibrate continuously. Simulate with confidence.

Get Started

  1. Open PraxiRun
  2. Go to Connectors → connect your Datadog/CloudWatch
  3. Click "Import Production Data" in the dashboard
  4. The simulation auto-calibrates
  5. Run what-if scenarios against your real baseline

Your architecture decisions are now backed by data, not debate.

Ready to test your architecture skills?

Try a Free Simulation →

Comments

No comments yet. Be the first!