Service Level Objectives (SLOs), SLAs & Observability Architecture
This living technical specification defines the reliability standards, observability infrastructure, health-checking protocols, and incident escalation frameworks for the Debelu campus marketplace. Grounded directly in [HealthCheckService.ts](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-backend/src/services/HealthCheckService.ts), [serviceStatus.ts](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-backend/src/lib/serviceStatus.ts), [logger.ts](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-backend/src/lib/logger.ts), and the perimeter routing in [server.ts](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-backend/src/server.ts), this document establishes contractual service guarantees, measurable telemetry signals, and automated operational thresholds.
1. Overview & Service Level Agreements (SLAs)
Debelu separates internal engineering objectives (SLOs) from legally and contractually binding commitments (SLAs) made to university students, campus vendors, university administrations, and institutional payment partners.
flowchart TD
subgraph Commitments [Service Level Framework]
SLA[Contractual SLA<br/>Business & Legal Commitments]
SLO[Internal Engineering SLO<br/>Strict Target Thresholds]
SLI[Service Level Indicators SLIs<br/>Real-Time Telemetry Metrics]
SLI -->|Evaluates Performance| SLO
SLO -->|Provides Safety Margin To| SLA
end
subgraph ErrorBudgetManagement [Burn Rate Control]
SLO -->|Downtime Allowance| EB[Error Budget: 100% - SLO]
EB -->|Rapid Burn: 14.4x| P1[SEV-1 Critical Page]
EB -->|Slow Burn: 6.0x| P2[SEV-2 Ticket / Warning]
EB -->|Exhaustion: 0% Left| Freeze[Automated Feature Freeze]
end1.1 Commercial & Operational SLA Guarantees
| Category | SLA Commitment | Target | Remedies / Penalty for Breach |
|---|---|---|---|
| Platform Availability | Monthly aggregate uptime of buyer storefront and checkout gateway. | $\ge 99.9%$ (< 43.2m downtime/mo) | Pro-rated commission credits to certified campus vendors on monthly platform fees. |
| Financial Escrow Release | Payout transfer initiation following counterparty PIN delivery verification. | $< 24$ hours | Automatic escalation to Tier-1 finance operations with priority disbursement queueing. |
| Dispute Arbitration | Time to triage and issue initial finding for an escrow dispute ticket. | $< 48$ hours | Buyer full refund issued automatically if vendor fails to respond to evidence requests within SLA. |
| Support First Response | Dedicated ticket response time for transaction/payment related inquiries. | $< 30$ mins (Critical) $< 2$ hours (General) | Automated escalation to Head of Operations; priority ticket badge in customer UI. |
| Data Subject Requests (NDPA) | Fulfillment of personal data portability exports and erasure verification. | $< 30$ days | Statutory compliance reporting to NDPC; immediate supervisory alert. |
2. Service Level Objectives (SLOs) & Indicators (SLIs) Matrix
Debelu enforces strict, measurable SLIs across all core customer journeys. Targets are evaluated over continuous rolling windows.
| Journey / Subsystem | Service Level Indicator (SLI) | Target (SLO) | Window | Telemetry Query / Metric Expression |
|---|---|---|---|---|
| Checkout & Payments | $\frac{\text{Successful 2xx Checkouts}}{\text{Total Checkout Payment Intents}}$ | $\ge 99.95%$ | Rolling 30 Days | sum(rate(http_requests_total{path=~"/api/payments.*",status=~"2.."}[5m])) / sum(rate(http_requests_total{path=~"/api/payments.*"}[5m])) |
| Checkout Latency | HTTP Roundtrip latency for payment intent creation (p95) | $\le 600$ ms | Rolling 7 Days | histogram_quantile(0.95, sum(rate(http_request_duration_ms_bucket{path="/api/payments/intent"}[5m])) by (le)) |
| Product Search & Catalog | Latency for product query and filtering endpoints (p95) | $\le 250$ ms | Rolling 7 Days | histogram_quantile(0.95, sum(rate(http_request_duration_ms_bucket{path="/api/products"}[5m])) by (le)) |
| Escrow State Transitions | $\frac{\text{Atomic Escrow Transitions Without Deadlock}}{\text{Total Escrow Release / Refund Invocations}}$ | $\ge 99.99%$ | Rolling 30 Days | sum(rate(escrow_transitions_success_total[5m])) / sum(rate(escrow_transitions_total[5m])) |
| Nduzi AI Assistance | Streaming response initiation time (Time To First Token) | $\le 800$ ms (p90) | Rolling 24 Hours | histogram_quantile(0.90, sum(rate(nduzi_ttft_ms_bucket[5m])) by (le)) |
| Webhook Delivery & Ingestion | Webhook ingestion latency from gateway to BullMQ queuing | $\le 200$ ms (p99) | Rolling 24 Hours | histogram_quantile(0.99, sum(rate(webhook_ingest_duration_ms_bucket[5m])) by (le)) |
| API Perimeter Health | HTTP 5xx Server Error Rate across all Express endpoints | $\le 0.05%$ | Rolling 24 Hours | sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) |
| Storefront Web Vitals | Core Web Vitals: Largest Contentful Paint (LCP) | $\le 2.2$ s (p75) | Rolling 7 Days | quantile_over_time(0.75, web_vitals_lcp_seconds[7d]) |
| Storefront Interaction | Core Web Vitals: Interaction to Next Paint (INP) | $\le 150$ ms (p75) | Rolling 7 Days | quantile_over_time(0.75, web_vitals_inp_ms[7d]) |
3. Error Budget & Multi-Window Multi-Burn Rate Alerting
Debelu adopts the Google Cloud SRE Multi-Window Multi-Burn Rate alerting model to eliminate alert fatigue while providing instantaneous paging for catastrophic failures.
3.1 Error Budget Calculations
The Error Budget is calculated as: $$\text{Error Budget} = 1.0 - \text{SLO}$$
For Debelu's core payment and order services ($\text{SLO} = 99.95%$ over 30 days):
- Total Requests: Assuming $1,000,000$ transactions/month, the error budget permits exactly 500 failed requests.
- Allowable Cumulative Downtime:
| Time Window | 99.9% Uptime (Public SLA) | 99.95% Uptime (Internal Checkout SLO) | 99.99% Uptime (Escrow Ledger SLO) |
|---|---|---|---|
| 1 Hour | 3.6 seconds | 1.8 seconds | 0.36 seconds |
| 24 Hours (1 Day) | 1.44 minutes | 43.2 seconds | 8.64 seconds |
| 7 Days (1 Week) | 10.08 minutes | 5.04 minutes | 1.01 minutes |
| 30 Days (1 Month) | 43.2 minutes | 21.6 minutes | 4.32 minutes |
| 365 Days (1 Year) | 8.76 hours | 4.38 hours | 52.56 minutes |
3.2 Burn Rate Alerting Matrix
Burn rate represents the consumption speed of the error budget. A burn rate of $1.0$ consumes $100%$ of the error budget over exactly 30 days.
Burn Rate (B) = (% Error Budget Consumed) / (% Period Elapsed)Debelu employs dual-window verification (short window to confirm current conditions, long window to avoid false positives) before triggering on-call pages:
| Severity Level | Burn Rate ($B$) | % Budget Consumed | Long Window (Alert Window) | Short Window (Reset Window) | Notification Target | Action Required |
|---|---|---|---|---|---|---|
| SEV-1 (Critical Page) | 14.4x | 2.0% | 1 hour | 5 minutes | PagerDuty Call to Primary On-Call | Immediate incident war room; trigger failover runbook. |
| SEV-1 (Catastrophic) | 14.4x | 5.0% | 6 hours | 30 minutes | PagerDuty to Escalation Team | Executive alert; initiate secondary gateway diversion. |
| SEV-2 (High Warning) | 6.0x | 5.0% | 6 hours | 30 minutes | Slack #ops-high-priority + SMS | On-call engineer triage within 30 minutes. |
| SEV-3 (Moderate Burn) | 1.0x | 10.0% | 3 days | 6 hours | Jira/Linear Bug Creation | Addressed in current sprint backlog. |
3.3 Automated Feature Freeze & Code Red Protocol
When an SLI violates its budget, the system executes an automated freeze policy:
stateDiagram-v2
[*] --> Green: Budget > 25%
Green --> Yellow: 10% <= Budget <= 25%
Yellow --> Green: Budget Recovers > 25%
Yellow --> Red: Budget < 10%
Red --> CodeRed: Budget <= 0% (Exhausted)
state Yellow {
[*] --> DiscretionarySlowdown
DiscretionarySlowdown --> FlagLowPriorityPRs
}
state Red {
[*] --> StrictPRApproval
StrictPRApproval --> MandatoryReliabilityTesting
}
state CodeRed {
[*] --> AutomatedPipelineLock
AutomatedPipelineLock --> FeatureDeploymentsBlocked
FeatureDeploymentsBlocked --> HotfixesOnly
}- Yellow State ($10% \le \text{Budget} \le 25%$ remaining):
- CI/CD emits warnings on non-critical feature pull requests.
- Engineering managers review scheduled releases for potential latency impacts.
- Red State ($0% < \text{Budget} < 10%$ remaining):
- Non-critical pull requests require dual-authorization from Principal Architect and QA Lead.
- Architectural and refactoring PRs without direct reliability improvements are deferred.
- Code Red State ($\text{Budget} \le 0%$):
- Automated GitHub Actions deployment pipeline locks feature merges.
- Only
fix/*andhotfix/*branches resolving active incident post-mortems or P1 security patches can be released to production. - 100% of engineering bandwidth is redirected to infrastructure stabilization until a 7-day rolling window restores budget above $25%$.
4. Health Check Engine & Diagnostics Pipeline
Debelu implements a split health-checking topology conforming to cloud-native orchestrator standards (Kubernetes, Fly.io, Railway, Google Cloud Run) and high-availability load balancers.
graph TD
LB[Load Balancer / Cloudflare] -->|Every 5s| Liveness["GET /health/live<br/>HTTP 200 Fast Return"]
LB -->|Every 10s| Readiness["GET /health/ready<br/>Postgres & Redis Check"]
Admin[SRE Dashboard / Ops Monitoring] -->|Every 30s| Diagnostics["GET /health<br/>HealthCheckService.checkAll()"]
MobileApp[Storefront / Mobile PWA] -->|Cached 30s| PublicStatus["GET /api/status<br/>Maintenance & Components"]
subgraph ServiceChecks [HealthCheckService Deep Probes]
Diagnostics --> ProbeDB[(Supabase Postgres)]
Diagnostics --> ProbePaystack[Paystack Gateway]
Diagnostics --> ProbeR2[(Cloudflare R2 Bucket)]
Diagnostics --> ProbeGemini[Gemini AI Models]
Diagnostics --> ProbeTermii[Termii SMS Gateway]
Diagnostics --> ProbeRedis[(Redis Cluster)]
Diagnostics --> ProbeQueue[BullMQ Webhook Queue]
end4.1 Probe Specifications
A. Public Liveness Probe (GET /health/live, HEAD /health/live)
- Purpose: Verifies that the Node.js event loop is unblocked and the Express process is capable of receiving socket connections.
- Payload:json
{ "status": "live" } - Performance Guarantee: Response latency $< 5$ ms. Does not execute asynchronous I/O or database queries.
B. Deep Readiness Probe (GET /health/ready, HEAD /health/ready)
- Purpose: Verifies that the pod or instance is fully prepared to handle live user traffic. If downstream infrastructure fails, load balancers immediately take the node out of service.
- Validation Logic (from
server.ts):- Executes
supabaseAdmin.from('platform_settings').select('id').limit(1). - If
process.env.REDIS_URLis configured, executesredis.ping().
- Executes
- Response:
- Healthy (200 OK):
{ "status": "ready" } - Degraded/Unhealthy (503 Service Unavailable):
{ "status": "not_ready" }
- Healthy (200 OK):
C. Comprehensive SRE Diagnostic Dashboard (GET /health)
- Purpose: Evaluates all subsystems and dependencies in real-time, providing immediate observability to SRE dashboards and incident commanders.
- Implementation:typescript
// debelu-backend/src/server.ts:198-241 app.get('/health', async (req, res) => { // Collects checks from redis, supabase, gemini, paystack, and BullMQ queues const allOk = (checks.redis === 'ok' || !process.env.REDIS_URL) && checks.supabase === 'ok' && checks.gemini === 'operational' && checks.paystack === 'operational'; res.status(200).json({ status: allOk ? 'ok' : 'degraded', message: 'Debelu Backend is running', version: process.env.GIT_SHA || process.env.RAILWAY_GIT_COMMIT_SHA || 'unknown', uptime: process.uptime(), timestamp: new Date().toISOString(), requestId: (req as any).requestId, checks }); });
D. Public Member Status Screen (GET /api/status)
- Purpose: Feeds the buyer storefront, vendor portal, and mobile apps with non-sensitive platform status and maintenance messaging.
- Caching: Enforces a strict 30-second memory cache (
TTL_MS = 30_000via [serviceStatus.ts](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-backend/src/lib/serviceStatus.ts)) to shield downstream databases from thundering-herd traffic during public outages.
4.2 The 7-Point Dependency Engine (HealthCheckService.ts)
Debelu's automated diagnostic engine runs seven independent, timeout-bounded dependency probes:
// debelu-backend/src/services/HealthCheckService.ts
export class HealthCheckService {
static async checkAll(): Promise<DependencyHealthCheck[]> {
return Promise.all([
this.database(),
this.paystack(),
this.storage(),
this.gemini(),
this.whatsapp(),
this.termii(),
this.cache()
]);
}
}classDiagram
class DependencyHealthCheck {
+Service service
+String name
+Status status
+Number latencyMs
+String lastCheckedAt
+String errorMessage
+Object details
}
class HealthCheckService {
+checkAll() Promise~DependencyHealthCheck[]~
-database()
-paystack()
-storage()
-gemini()
-whatsapp()
-termii()
-cache()
}
HealthCheckService ..> DependencyHealthCheck : producesDetailed Probe Breakdown
| Dependency Probe | Target & Method | Timeout | Degradation Threshold | Failure Signature & Classification |
|---|---|---|---|---|
supabase | supabaseAdmin.from('admin_roles').select('id').limit(1) | 4,000 ms | Latency $> 1,200$ ms | status: "down" on timeout or SQL connection reset. |
paystack | FinancialObservationService.gateway() (/balance) | 5,000 ms | Test mode active or latency $> 1,500$ ms | status: "unavailable" if API key unconfigured or invalid NUBAN switch. |
r2_storage | @aws-sdk/client-s3 HeadBucketCommand to Cloudflare R2 | 4,000 ms | Latency $> 1,500$ ms | status: "down" if R2 credentials invalid or bucket unverified. |
gemini | GET /v1beta/models/{model} with x-goog-api-key | 5,000 ms | Latency $> 1,500$ ms | status: "unavailable" if model missing; status: "down" on timeout. |
whatsapp | Meta Graph API token configuration and phone ID check | N/A | Provider unprobed | status: "unavailable" with credentialsConfigured boolean detail. |
termii_sms | GET /api/get-balance with Termii API Key | 4,000 ms | Latency $> 1,500$ ms | status: "down" if 5xx HTTP response; validates NGN currency format. |
cache | redis.ping() raced against fallback timeout | 4,000 ms | Latency $> 1,200$ ms | status: "down" if socket disconnected or ping returns non-PONG. |
5. Telemetry, Structured Logging & PII Scrubbing
Debelu employs centralized structured logging conforming to 12-factor application design, mandating end-to-end trace correlation and cryptographic data redacting.
5.1 Correlation Protocol & Request Tracing
Every inbound HTTP request receives an immutable x-request-id assigned at the perimeter middleware. If the client or upstream proxy supplies a valid UUID v4 x-request-id, it is preserved; otherwise, a fresh UUID v4 is generated.
// debelu-backend/src/lib/logger.ts
export const logWithRequest = (
req: any,
level: 'info' | 'warn' | 'error' | 'debug',
msg: string,
meta: any = {}
) => {
logger.log(level, msg, {
requestId: req.requestId,
path: req.path,
method: req.method,
campusScope: req.user?.campusScope,
...meta
});
};5.2 Structured JSON Log Schema
All stdout emissions in production follow strict JSON formatting:
{
"timestamp": "2026-10-05 09:58:12",
"level": "info",
"message": "Payment intent created successfully",
"service": "debelu-backend",
"version": "f48c1b9",
"requestId": "d8e3b092-7f94-4d8b-967f-4421d0f507ba",
"path": "/api/payments/intent",
"method": "POST",
"orderId": "ord_88201491",
"amountKobo": 450000,
"campus": "UNILAG",
"durationMs": 142
}5.3 Automated PII & Credential Redaction
Before any log object is serialized to stdout or transported to log drains, it passes through the recursive sanitizer (redactValue from redaction.ts):
// debelu-backend/src/lib/logger.ts:13-20
format: combine(
winston.format((info) => {
for (const key of Object.keys(info)) {
const clean = redactValue({ [key]: info[key] }) as Record<string, unknown>;
info[key] = clean[key];
}
return info;
})(),
timestamp({ format: 'YYYY-MM-DD HH:mm:ss' }),
process.env.NODE_ENV === 'production' ? json() : combine(colorize(), myFormat)
)Mandatory Redacted Fields:
- Passwords and PIN codes (
delivery_pin,password,pin_hash) - Financial secrets (
PAYSTACK_SECRET_KEY,authorization_code,card_last4[isolated]) - PII and Identity credentials (
bvn,nin,account_number,phone_number) - Bearer tokens, JWTs, and Supabase service keys (
Bearer eyJ...replaced with[REDACTED])
5.4 Client & Server Error Tracking (Sentry)
- Backend API: Captures unhandled promise rejections, Express error middleware catches, and database serialization failures (
40001). - Storefront PWA: Wraps React component trees in custom Error Boundaries (
StorefrontErrorBoundary), capturing client crashes, network disconnects, and offline caching errors. - Scrubbing Rule: Sentry SDKs configure
beforeSendhooks to scrub transaction query strings, authorization headers, and form input bodies before transmission.
6. On-Call Escalation Matrix & Incident Management
Debelu operates an automated Incident Command System (ICS) to manage critical failures, operational outages, and data security incidents.
sequenceDiagram
autonumber
participant Alert as Monitoring / Alertmanager
participant Primary as Level 1: Primary On-Call
participant Secondary as Level 2: Backup On-Call
participant Exec as Level 3: Exec & Principal SRE
Alert->>Primary: SEV-1 Triggered (PagerDuty Call)
Note over Primary: ACK Required Within 15 Mins
alt Primary Acknowledges
Primary->>Alert: Acknowledged & Opens War Room
else 15 Mins Unacknowledged
Alert->>Secondary: Escalates to Secondary On-Call (SMS + Call)
Note over Secondary: ACK Required Within 10 Mins
alt Secondary Acknowledges
Secondary->>Alert: Acknowledged & Opens War Room
else 25 Mins Total Unacknowledged
Alert->>Exec: PagerDuty Page to Head of Engineering & CTO
Exec->>Alert: Emergency Command Assumed
end
end6.1 Severity Classification Scheme
| Severity | Operational Definition | Example Incidents | Initial Response SLA | Status Page Communication |
|---|---|---|---|---|
| SEV-1 (Critical) | Core business failure; payment processing halted, platform down, or data security breach. | Paystack gateway 502, Supabase database inaccessible, active card/escrow exploit. | $< 15$ mins | Public incident posted within 15 mins; updates every 30 mins. |
| SEV-2 (Major) | Major component degraded; workaround exists but core UX impaired. | Nduzi AI assistant offline, vendor file uploads failing, search response $> 1.5$s. | $< 30$ mins | Component marked degraded on /api/status; updates every 60 mins. |
| SEV-3 (Moderate) | Minor feature bug, internal tool failure, or non-blocking performance degradation. | Admin export failure, single campus banner missing, delayed non-critical email. | $< 4$ hours | Internal ops notice; no public status page update required. |
| SEV-4 (Low) | Cosmetic UI defects, documentation typo, non-urgent feature question. | Mobile storefront styling flaw, minor layout shift on tablet. | Next Business Day | Internal tracking via Linear/GitHub issues. |
6.2 Escalation Roster & Responsibilities
- Level 1 — Primary On-Call Engineer:
- Rotates weekly across full-stack and backend engineering teams.
- Responsible for acknowledging pages within 15 minutes, triaging root cause, opening the incident Google Meet / Slack war room, and initiating documented runbooks.
- Level 2 — Secondary On-Call Engineer:
- Senior SRE or backend specialist.
- Paged automatically if the primary fails to acknowledge within 15 minutes or if the primary requests specialist escalation for database replication or payment reconciliation failures.
- Level 3 — Incident Commander & Principal Architect:
- Paged if a SEV-1 remains active and unresolved after 45 minutes.
- Authorized to declare emergency disaster recovery failovers, execute platform maintenance mode, or order DNS/Cloudflare rerouting.
- Executive Communications Lead (CTO / Head of Product):
- Manages communication with institutional partners, university student unions, media, and data protection regulatory authorities (NDPC / CBN).
6.3 Post-Mortem & Continuous Reliability Review
All SEV-1 and SEV-2 incidents mandate a blameless post-mortem completed within 72 hours of incident resolution:
- Timeline of Events: Chronological log of detection, paging, triage steps, and resolution.
- Root Cause Analysis (5 Whys): Deep technical exploration identifying procedural and systemic vulnerabilities rather than human error.
- Corrective & Preventive Actions (CAPA): Action items assigned to specific engineers with mandatory completion SLAs (P1 action items $\le 14$ days, P2 action items $\le 30$ days).
- Error Budget Restitution: Impact assessment quantifying total error budget consumed and determining whether feature deployments must be paused.
7. Document Revision History
| Revision | Date | Author / SRE Role | Scope of Changes | Status |
|---|---|---|---|---|
1.0.0 | 2026-10-05 | Lead SRE & Reliability Operations | Initial enterprise living specification for SLOs, SLAs, multi-burn rate alerting, HealthCheckService 7-point telemetry, and incident response matrices. | Active Living Standard |