Skip to content

Service Level Objectives (SLOs), SLAs & Observability Architecture ​

This living technical specification defines the reliability standards, observability infrastructure, health-checking protocols, and incident escalation frameworks for the Debelu campus marketplace. Grounded directly in [HealthCheckService.ts](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-backend/src/services/HealthCheckService.ts), [serviceStatus.ts](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-backend/src/lib/serviceStatus.ts), [logger.ts](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-backend/src/lib/logger.ts), and the perimeter routing in [server.ts](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-backend/src/server.ts), this document establishes contractual service guarantees, measurable telemetry signals, and automated operational thresholds.


1. Overview & Service Level Agreements (SLAs) ​

Debelu separates internal engineering objectives (SLOs) from legally and contractually binding commitments (SLAs) made to university students, campus vendors, university administrations, and institutional payment partners.

mermaid
flowchart TD
    subgraph Commitments [Service Level Framework]
        SLA[Contractual SLA<br/>Business & Legal Commitments]
        SLO[Internal Engineering SLO<br/>Strict Target Thresholds]
        SLI[Service Level Indicators SLIs<br/>Real-Time Telemetry Metrics]
        
        SLI -->|Evaluates Performance| SLO
        SLO -->|Provides Safety Margin To| SLA
    end

    subgraph ErrorBudgetManagement [Burn Rate Control]
        SLO -->|Downtime Allowance| EB[Error Budget: 100% - SLO]
        EB -->|Rapid Burn: 14.4x| P1[SEV-1 Critical Page]
        EB -->|Slow Burn: 6.0x| P2[SEV-2 Ticket / Warning]
        EB -->|Exhaustion: 0% Left| Freeze[Automated Feature Freeze]
    end

1.1 Commercial & Operational SLA Guarantees ​

CategorySLA CommitmentTargetRemedies / Penalty for Breach
Platform AvailabilityMonthly aggregate uptime of buyer storefront and checkout gateway.$\ge 99.9%$ (< 43.2m downtime/mo)Pro-rated commission credits to certified campus vendors on monthly platform fees.
Financial Escrow ReleasePayout transfer initiation following counterparty PIN delivery verification.$< 24$ hoursAutomatic escalation to Tier-1 finance operations with priority disbursement queueing.
Dispute ArbitrationTime to triage and issue initial finding for an escrow dispute ticket.$< 48$ hoursBuyer full refund issued automatically if vendor fails to respond to evidence requests within SLA.
Support First ResponseDedicated ticket response time for transaction/payment related inquiries.$< 30$ mins (Critical)
$< 2$ hours (General)
Automated escalation to Head of Operations; priority ticket badge in customer UI.
Data Subject Requests (NDPA)Fulfillment of personal data portability exports and erasure verification.$< 30$ daysStatutory compliance reporting to NDPC; immediate supervisory alert.

2. Service Level Objectives (SLOs) & Indicators (SLIs) Matrix ​

Debelu enforces strict, measurable SLIs across all core customer journeys. Targets are evaluated over continuous rolling windows.

Journey / SubsystemService Level Indicator (SLI)Target (SLO)WindowTelemetry Query / Metric Expression
Checkout & Payments$\frac{\text{Successful 2xx Checkouts}}{\text{Total Checkout Payment Intents}}$$\ge 99.95%$Rolling 30 Dayssum(rate(http_requests_total{path=~"/api/payments.*",status=~"2.."}[5m])) / sum(rate(http_requests_total{path=~"/api/payments.*"}[5m]))
Checkout LatencyHTTP Roundtrip latency for payment intent creation (p95)$\le 600$ msRolling 7 Dayshistogram_quantile(0.95, sum(rate(http_request_duration_ms_bucket{path="/api/payments/intent"}[5m])) by (le))
Product Search & CatalogLatency for product query and filtering endpoints (p95)$\le 250$ msRolling 7 Dayshistogram_quantile(0.95, sum(rate(http_request_duration_ms_bucket{path="/api/products"}[5m])) by (le))
Escrow State Transitions$\frac{\text{Atomic Escrow Transitions Without Deadlock}}{\text{Total Escrow Release / Refund Invocations}}$$\ge 99.99%$Rolling 30 Dayssum(rate(escrow_transitions_success_total[5m])) / sum(rate(escrow_transitions_total[5m]))
Nduzi AI AssistanceStreaming response initiation time (Time To First Token)$\le 800$ ms (p90)Rolling 24 Hourshistogram_quantile(0.90, sum(rate(nduzi_ttft_ms_bucket[5m])) by (le))
Webhook Delivery & IngestionWebhook ingestion latency from gateway to BullMQ queuing$\le 200$ ms (p99)Rolling 24 Hourshistogram_quantile(0.99, sum(rate(webhook_ingest_duration_ms_bucket[5m])) by (le))
API Perimeter HealthHTTP 5xx Server Error Rate across all Express endpoints$\le 0.05%$Rolling 24 Hourssum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
Storefront Web VitalsCore Web Vitals: Largest Contentful Paint (LCP)$\le 2.2$ s (p75)Rolling 7 Daysquantile_over_time(0.75, web_vitals_lcp_seconds[7d])
Storefront InteractionCore Web Vitals: Interaction to Next Paint (INP)$\le 150$ ms (p75)Rolling 7 Daysquantile_over_time(0.75, web_vitals_inp_ms[7d])

3. Error Budget & Multi-Window Multi-Burn Rate Alerting ​

Debelu adopts the Google Cloud SRE Multi-Window Multi-Burn Rate alerting model to eliminate alert fatigue while providing instantaneous paging for catastrophic failures.

3.1 Error Budget Calculations ​

The Error Budget is calculated as: $$\text{Error Budget} = 1.0 - \text{SLO}$$

For Debelu's core payment and order services ($\text{SLO} = 99.95%$ over 30 days):

  • Total Requests: Assuming $1,000,000$ transactions/month, the error budget permits exactly 500 failed requests.
  • Allowable Cumulative Downtime:
Time Window99.9% Uptime (Public SLA)99.95% Uptime (Internal Checkout SLO)99.99% Uptime (Escrow Ledger SLO)
1 Hour3.6 seconds1.8 seconds0.36 seconds
24 Hours (1 Day)1.44 minutes43.2 seconds8.64 seconds
7 Days (1 Week)10.08 minutes5.04 minutes1.01 minutes
30 Days (1 Month)43.2 minutes21.6 minutes4.32 minutes
365 Days (1 Year)8.76 hours4.38 hours52.56 minutes

3.2 Burn Rate Alerting Matrix ​

Burn rate represents the consumption speed of the error budget. A burn rate of $1.0$ consumes $100%$ of the error budget over exactly 30 days.

Burn Rate (B) = (% Error Budget Consumed) / (% Period Elapsed)

Debelu employs dual-window verification (short window to confirm current conditions, long window to avoid false positives) before triggering on-call pages:

Severity LevelBurn Rate ($B$)% Budget ConsumedLong Window (Alert Window)Short Window (Reset Window)Notification TargetAction Required
SEV-1 (Critical Page)14.4x2.0%1 hour5 minutesPagerDuty Call to Primary On-CallImmediate incident war room; trigger failover runbook.
SEV-1 (Catastrophic)14.4x5.0%6 hours30 minutesPagerDuty to Escalation TeamExecutive alert; initiate secondary gateway diversion.
SEV-2 (High Warning)6.0x5.0%6 hours30 minutesSlack #ops-high-priority + SMSOn-call engineer triage within 30 minutes.
SEV-3 (Moderate Burn)1.0x10.0%3 days6 hoursJira/Linear Bug CreationAddressed in current sprint backlog.

3.3 Automated Feature Freeze & Code Red Protocol ​

When an SLI violates its budget, the system executes an automated freeze policy:

mermaid
stateDiagram-v2
    [*] --> Green: Budget > 25%
    Green --> Yellow: 10% <= Budget <= 25%
    Yellow --> Green: Budget Recovers > 25%
    Yellow --> Red: Budget < 10%
    Red --> CodeRed: Budget <= 0% (Exhausted)
    
    state Yellow {
        [*] --> DiscretionarySlowdown
        DiscretionarySlowdown --> FlagLowPriorityPRs
    }
    
    state Red {
        [*] --> StrictPRApproval
        StrictPRApproval --> MandatoryReliabilityTesting
    }
    
    state CodeRed {
        [*] --> AutomatedPipelineLock
        AutomatedPipelineLock --> FeatureDeploymentsBlocked
        FeatureDeploymentsBlocked --> HotfixesOnly
    }
  1. Yellow State ($10% \le \text{Budget} \le 25%$ remaining):
    • CI/CD emits warnings on non-critical feature pull requests.
    • Engineering managers review scheduled releases for potential latency impacts.
  2. Red State ($0% < \text{Budget} < 10%$ remaining):
    • Non-critical pull requests require dual-authorization from Principal Architect and QA Lead.
    • Architectural and refactoring PRs without direct reliability improvements are deferred.
  3. Code Red State ($\text{Budget} \le 0%$):
    • Automated GitHub Actions deployment pipeline locks feature merges.
    • Only fix/* and hotfix/* branches resolving active incident post-mortems or P1 security patches can be released to production.
    • 100% of engineering bandwidth is redirected to infrastructure stabilization until a 7-day rolling window restores budget above $25%$.

4. Health Check Engine & Diagnostics Pipeline ​

Debelu implements a split health-checking topology conforming to cloud-native orchestrator standards (Kubernetes, Fly.io, Railway, Google Cloud Run) and high-availability load balancers.

mermaid
graph TD
    LB[Load Balancer / Cloudflare] -->|Every 5s| Liveness["GET /health/live<br/>HTTP 200 Fast Return"]
    LB -->|Every 10s| Readiness["GET /health/ready<br/>Postgres & Redis Check"]
    Admin[SRE Dashboard / Ops Monitoring] -->|Every 30s| Diagnostics["GET /health<br/>HealthCheckService.checkAll()"]
    MobileApp[Storefront / Mobile PWA] -->|Cached 30s| PublicStatus["GET /api/status<br/>Maintenance & Components"]

    subgraph ServiceChecks [HealthCheckService Deep Probes]
        Diagnostics --> ProbeDB[(Supabase Postgres)]
        Diagnostics --> ProbePaystack[Paystack Gateway]
        Diagnostics --> ProbeR2[(Cloudflare R2 Bucket)]
        Diagnostics --> ProbeGemini[Gemini AI Models]
        Diagnostics --> ProbeTermii[Termii SMS Gateway]
        Diagnostics --> ProbeRedis[(Redis Cluster)]
        Diagnostics --> ProbeQueue[BullMQ Webhook Queue]
    end

4.1 Probe Specifications ​

A. Public Liveness Probe (GET /health/live, HEAD /health/live) ​

  • Purpose: Verifies that the Node.js event loop is unblocked and the Express process is capable of receiving socket connections.
  • Payload:
    json
    { "status": "live" }
  • Performance Guarantee: Response latency $< 5$ ms. Does not execute asynchronous I/O or database queries.

B. Deep Readiness Probe (GET /health/ready, HEAD /health/ready) ​

  • Purpose: Verifies that the pod or instance is fully prepared to handle live user traffic. If downstream infrastructure fails, load balancers immediately take the node out of service.
  • Validation Logic (from server.ts):
    1. Executes supabaseAdmin.from('platform_settings').select('id').limit(1).
    2. If process.env.REDIS_URL is configured, executes redis.ping().
  • Response:
    • Healthy (200 OK): { "status": "ready" }
    • Degraded/Unhealthy (503 Service Unavailable): { "status": "not_ready" }

C. Comprehensive SRE Diagnostic Dashboard (GET /health) ​

  • Purpose: Evaluates all subsystems and dependencies in real-time, providing immediate observability to SRE dashboards and incident commanders.
  • Implementation:
    typescript
    // debelu-backend/src/server.ts:198-241
    app.get('/health', async (req, res) => {
      // Collects checks from redis, supabase, gemini, paystack, and BullMQ queues
      const allOk = (checks.redis === 'ok' || !process.env.REDIS_URL) &&
                    checks.supabase === 'ok' &&
                    checks.gemini === 'operational' &&
                    checks.paystack === 'operational';
      res.status(200).json({
        status: allOk ? 'ok' : 'degraded',
        message: 'Debelu Backend is running',
        version: process.env.GIT_SHA || process.env.RAILWAY_GIT_COMMIT_SHA || 'unknown',
        uptime: process.uptime(),
        timestamp: new Date().toISOString(),
        requestId: (req as any).requestId,
        checks
      });
    });

D. Public Member Status Screen (GET /api/status) ​

  • Purpose: Feeds the buyer storefront, vendor portal, and mobile apps with non-sensitive platform status and maintenance messaging.
  • Caching: Enforces a strict 30-second memory cache (TTL_MS = 30_000 via [serviceStatus.ts](file:///c:/Users/frank/OneDrive/Desktop/Chisom/Debelu/New%20Debelu%20Marketplace/debelu-backend/src/lib/serviceStatus.ts)) to shield downstream databases from thundering-herd traffic during public outages.

4.2 The 7-Point Dependency Engine (HealthCheckService.ts) ​

Debelu's automated diagnostic engine runs seven independent, timeout-bounded dependency probes:

typescript
// debelu-backend/src/services/HealthCheckService.ts
export class HealthCheckService {
    static async checkAll(): Promise<DependencyHealthCheck[]> {
        return Promise.all([
            this.database(),
            this.paystack(),
            this.storage(),
            this.gemini(),
            this.whatsapp(),
            this.termii(),
            this.cache()
        ]);
    }
}
mermaid
classDiagram
    class DependencyHealthCheck {
        +Service service
        +String name
        +Status status
        +Number latencyMs
        +String lastCheckedAt
        +String errorMessage
        +Object details
    }
    class HealthCheckService {
        +checkAll() Promise~DependencyHealthCheck[]~
        -database()
        -paystack()
        -storage()
        -gemini()
        -whatsapp()
        -termii()
        -cache()
    }
    HealthCheckService ..> DependencyHealthCheck : produces

Detailed Probe Breakdown ​

Dependency ProbeTarget & MethodTimeoutDegradation ThresholdFailure Signature & Classification
supabasesupabaseAdmin.from('admin_roles').select('id').limit(1)4,000 msLatency $> 1,200$ msstatus: "down" on timeout or SQL connection reset.
paystackFinancialObservationService.gateway() (/balance)5,000 msTest mode active or latency $> 1,500$ msstatus: "unavailable" if API key unconfigured or invalid NUBAN switch.
r2_storage@aws-sdk/client-s3 HeadBucketCommand to Cloudflare R24,000 msLatency $> 1,500$ msstatus: "down" if R2 credentials invalid or bucket unverified.
geminiGET /v1beta/models/{model} with x-goog-api-key5,000 msLatency $> 1,500$ msstatus: "unavailable" if model missing; status: "down" on timeout.
whatsappMeta Graph API token configuration and phone ID checkN/AProvider unprobedstatus: "unavailable" with credentialsConfigured boolean detail.
termii_smsGET /api/get-balance with Termii API Key4,000 msLatency $> 1,500$ msstatus: "down" if 5xx HTTP response; validates NGN currency format.
cacheredis.ping() raced against fallback timeout4,000 msLatency $> 1,200$ msstatus: "down" if socket disconnected or ping returns non-PONG.

5. Telemetry, Structured Logging & PII Scrubbing ​

Debelu employs centralized structured logging conforming to 12-factor application design, mandating end-to-end trace correlation and cryptographic data redacting.

5.1 Correlation Protocol & Request Tracing ​

Every inbound HTTP request receives an immutable x-request-id assigned at the perimeter middleware. If the client or upstream proxy supplies a valid UUID v4 x-request-id, it is preserved; otherwise, a fresh UUID v4 is generated.

typescript
// debelu-backend/src/lib/logger.ts
export const logWithRequest = (
    req: any,
    level: 'info' | 'warn' | 'error' | 'debug',
    msg: string,
    meta: any = {}
) => {
    logger.log(level, msg, {
        requestId: req.requestId,
        path: req.path,
        method: req.method,
        campusScope: req.user?.campusScope,
        ...meta
    });
};

5.2 Structured JSON Log Schema ​

All stdout emissions in production follow strict JSON formatting:

json
{
  "timestamp": "2026-10-05 09:58:12",
  "level": "info",
  "message": "Payment intent created successfully",
  "service": "debelu-backend",
  "version": "f48c1b9",
  "requestId": "d8e3b092-7f94-4d8b-967f-4421d0f507ba",
  "path": "/api/payments/intent",
  "method": "POST",
  "orderId": "ord_88201491",
  "amountKobo": 450000,
  "campus": "UNILAG",
  "durationMs": 142
}

5.3 Automated PII & Credential Redaction ​

Before any log object is serialized to stdout or transported to log drains, it passes through the recursive sanitizer (redactValue from redaction.ts):

typescript
// debelu-backend/src/lib/logger.ts:13-20
format: combine(
    winston.format((info) => {
        for (const key of Object.keys(info)) {
            const clean = redactValue({ [key]: info[key] }) as Record<string, unknown>;
            info[key] = clean[key];
        }
        return info;
    })(),
    timestamp({ format: 'YYYY-MM-DD HH:mm:ss' }),
    process.env.NODE_ENV === 'production' ? json() : combine(colorize(), myFormat)
)

Mandatory Redacted Fields:

  • Passwords and PIN codes (delivery_pin, password, pin_hash)
  • Financial secrets (PAYSTACK_SECRET_KEY, authorization_code, card_last4 [isolated])
  • PII and Identity credentials (bvn, nin, account_number, phone_number)
  • Bearer tokens, JWTs, and Supabase service keys (Bearer eyJ... replaced with [REDACTED])

5.4 Client & Server Error Tracking (Sentry) ​

  • Backend API: Captures unhandled promise rejections, Express error middleware catches, and database serialization failures (40001).
  • Storefront PWA: Wraps React component trees in custom Error Boundaries (StorefrontErrorBoundary), capturing client crashes, network disconnects, and offline caching errors.
  • Scrubbing Rule: Sentry SDKs configure beforeSend hooks to scrub transaction query strings, authorization headers, and form input bodies before transmission.

6. On-Call Escalation Matrix & Incident Management ​

Debelu operates an automated Incident Command System (ICS) to manage critical failures, operational outages, and data security incidents.

mermaid
sequenceDiagram
    autonumber
    participant Alert as Monitoring / Alertmanager
    participant Primary as Level 1: Primary On-Call
    participant Secondary as Level 2: Backup On-Call
    participant Exec as Level 3: Exec & Principal SRE
    
    Alert->>Primary: SEV-1 Triggered (PagerDuty Call)
    Note over Primary: ACK Required Within 15 Mins
    alt Primary Acknowledges
        Primary->>Alert: Acknowledged & Opens War Room
    else 15 Mins Unacknowledged
        Alert->>Secondary: Escalates to Secondary On-Call (SMS + Call)
        Note over Secondary: ACK Required Within 10 Mins
        alt Secondary Acknowledges
            Secondary->>Alert: Acknowledged & Opens War Room
        else 25 Mins Total Unacknowledged
            Alert->>Exec: PagerDuty Page to Head of Engineering & CTO
            Exec->>Alert: Emergency Command Assumed
        end
    end

6.1 Severity Classification Scheme ​

SeverityOperational DefinitionExample IncidentsInitial Response SLAStatus Page Communication
SEV-1 (Critical)Core business failure; payment processing halted, platform down, or data security breach.Paystack gateway 502, Supabase database inaccessible, active card/escrow exploit.$< 15$ minsPublic incident posted within 15 mins; updates every 30 mins.
SEV-2 (Major)Major component degraded; workaround exists but core UX impaired.Nduzi AI assistant offline, vendor file uploads failing, search response $> 1.5$s.$< 30$ minsComponent marked degraded on /api/status; updates every 60 mins.
SEV-3 (Moderate)Minor feature bug, internal tool failure, or non-blocking performance degradation.Admin export failure, single campus banner missing, delayed non-critical email.$< 4$ hoursInternal ops notice; no public status page update required.
SEV-4 (Low)Cosmetic UI defects, documentation typo, non-urgent feature question.Mobile storefront styling flaw, minor layout shift on tablet.Next Business DayInternal tracking via Linear/GitHub issues.

6.2 Escalation Roster & Responsibilities ​

  1. Level 1 — Primary On-Call Engineer:
    • Rotates weekly across full-stack and backend engineering teams.
    • Responsible for acknowledging pages within 15 minutes, triaging root cause, opening the incident Google Meet / Slack war room, and initiating documented runbooks.
  2. Level 2 — Secondary On-Call Engineer:
    • Senior SRE or backend specialist.
    • Paged automatically if the primary fails to acknowledge within 15 minutes or if the primary requests specialist escalation for database replication or payment reconciliation failures.
  3. Level 3 — Incident Commander & Principal Architect:
    • Paged if a SEV-1 remains active and unresolved after 45 minutes.
    • Authorized to declare emergency disaster recovery failovers, execute platform maintenance mode, or order DNS/Cloudflare rerouting.
  4. Executive Communications Lead (CTO / Head of Product):
    • Manages communication with institutional partners, university student unions, media, and data protection regulatory authorities (NDPC / CBN).

6.3 Post-Mortem & Continuous Reliability Review ​

All SEV-1 and SEV-2 incidents mandate a blameless post-mortem completed within 72 hours of incident resolution:

  1. Timeline of Events: Chronological log of detection, paging, triage steps, and resolution.
  2. Root Cause Analysis (5 Whys): Deep technical exploration identifying procedural and systemic vulnerabilities rather than human error.
  3. Corrective & Preventive Actions (CAPA): Action items assigned to specific engineers with mandatory completion SLAs (P1 action items $\le 14$ days, P2 action items $\le 30$ days).
  4. Error Budget Restitution: Impact assessment quantifying total error budget consumed and determining whether feature deployments must be paused.

7. Document Revision History ​

RevisionDateAuthor / SRE RoleScope of ChangesStatus
1.0.02026-10-05Lead SRE & Reliability OperationsInitial enterprise living specification for SLOs, SLAs, multi-burn rate alerting, HealthCheckService 7-point telemetry, and incident response matrices.Active Living Standard

Released under Proprietary Enterprise License.