Files
trustos/docs/ARCHITECTURE.md
drjones 9dbf59b995 docs: Add comprehensive documentation suite
- Enhance README.md with full installation instructions, architecture overview, features, configuration, testing, deployment, security, and troubleshooting sections
- Add ARCHITECTURE.md with detailed system architecture, data model, authentication, API design, frontend/backend architecture, database design, AI integration, security, and scalability considerations
-Add API.md with complete API reference including authentication, all endpoints, data models, examples, and interactive documentation links
- Add DEPLOYMENT.md with deployment guides for Railway, Render, VPS, and Kubernetes, including pre-deployment checklist, monitoring, backup, and troubleshooting
- Add CONTRIBUTING.md with development workflow, coding standards, testing guidelines, documentation standards, PR process, and community guidelines
2026-07-06 04:46:47 +00:00

35 KiB

TrustOS Architecture Documentation

This document provides a detailed overview of the TrustOS system architecture, design decisions, data flow, and technical implementation details.


Table of Contents


System Overview

TrustOS is a multi-tenant, AI-powered cyber resilience platform built on a modern microservices-inspired architecture. The system is designed to be:

  • Secure: Multi-tenant isolation with defense-in-depth security
  • Scalable: Async I/O throughout for high concurrency
  • Maintainable: Clean separation of concerns and modular design
  • Resilient: Graceful degradation when external services are unavailable

High-Level Architecture

┌─────────────────────────────────────────────────────────────────┐
│                        Client Layer                             │
│  Web Browser (Executive, IT Admin, TrustOS Admin)              │
└────────────────────┬────────────────────────────────────────────┘
                     │ HTTPS
┌────────────────────▼────────────────────────────────────────────┐
│                      Frontend Layer                             │
│  Next.js 16 + TypeScript + Tailwind CSS + shadcn/ui             │
│  - Server-Side Rendering (SSR)                                  │
│  - Client-Side Hydration                                         │
│  - Static Site Generation (SSG) where applicable                 │
└────────────────────┬────────────────────────────────────────────┘
                     │ REST API (JSON)
┌────────────────────▼────────────────────────────────────────────┐
│                       API Gateway                               │
│  FastAPI Application                                            │
│  - Request Validation (Pydantic)                                 │
│  - Authentication (JWT)                                          │
│  - Authorization (RBAC)                                          │
│  - Rate Limiting (future)                                       │
│  - Request Logging                                              │
└────────────────────┬────────────────────────────────────────────┘
                     │
        ┌────────────┴────────────┐
        │                         │
┌───────▼────────┐       ┌────────▼─────────┐
│  Service Layer │       │  Background      │
│                │       │  Workers         │
│ - Dashboard    │       │ - AI Translation│
│ - Findings     │       │ - Risk Calc     │
│ - Reports      │       │ - PDF Gen       │
│ - Footprint    │       │ - Monitoring    │
└───────┬────────┘       └────────┬─────────┘
        │                         │
┌───────▼─────────────────────────▼──────────┐
│              Data Access Layer                │
│  SQLAlchemy 2.0 (Async ORM)                   │
│  - Query Building                            │
│  - Connection Pooling                        │
│  - Transaction Management                   │
└────────────────────┬─────────────────────────┘
                     │
┌────────────────────▼─────────────────────────┐
│              Database Layer                   │
│  PostgreSQL 16                                │
│  - Multi-Tenant Data Isolation               │
│  - Indexing Strategy                         │
│  - Foreign Key Constraints                   │
│  - JSONB for Flexible Data                   │
└──────────────────────────────────────────────┘

External Services:
┌──────────────┐  ┌──────────────┐  ┌──────────────┐
│ OpenAI API   │  │ Anthropic API│  │ HIBP API     │
└──────────────┘  └──────────────┘  └──────────────┘

Architecture Principles

1. Authorization First

TrustOS only scans and monitors assets that have been explicitly authorized by the tenant organization. This is a core security and privacy principle:

  • Authorized Assets Table: All monitored assets must be pre-registered
  • Scope Enforcement: All automated checks respect the authorized asset list
  • Executive Enrollment: Executive monitoring requires explicit organizational consent
  • Audit Trail: All authorization decisions are logged

2. Multi-Tenant Isolation

Every data access path enforces tenant isolation at multiple layers:

  • Database Level: All tables include tenant_id with foreign key constraints
  • ORM Level: Queries automatically filter by tenant_id
  • API Level: Middleware validates tenant access before processing requests
  • Application Level: UI components only display tenant-specific data

3. Privacy by Design

Executive and organizational data is handled with privacy as a foundational requirement:

  • Minimal Data Collection: Only collect data necessary for security assessments
  • Explicit Consent: Executive enrollment requires organizational authorization
  • Data Minimization: Store only what, not who, where possible
  • Audit Logging: All data access is logged for accountability

4. AI-Augmented, Not AI-Dependent

AI features enhance the product but are not required for core functionality:

  • Graceful Degradation: Features work with raw technical data if AI is unavailable
  • Caching: AI responses are cached to avoid redundant API calls
  • Fallback Mechanisms: System continues operating if AI services are down
  • Cost Control: Rate limiting and caching to manage AI API costs

5. Audit Trail

All state changes are tracked with full provenance:

  • Timestamps: created_at and updated_at on all records
  • User Attribution: created_by and updated_by where applicable
  • Status Changes: Finding status transitions are logged with notes
  • Access Logs: API requests are logged with user and tenant context

Component Architecture

Frontend Components

Page Components

  • Dashboard Page (/dashboard): Executive view with risk score, top risks, trends
  • Findings Page (/findings): IT admin view with sortable/filterable table
  • Finding Detail Page (/findings/[id]): Detailed view with technical and business impact
  • Login Page (/login): Authentication interface
  • Footprint Page (/footprint): Digital footprint center

Reusable Components)

  • RiskDial: Circular gauge for cyber health score (0-100)
  • RiskCard: Card displaying finding with AI summary and impact
  • TrendChart: Line chart for 90-day risk history
  • RemediationBoard: Kanban board for finding status tracking
  • Badge: Severity and status badges with color coding

Hooks

  • useAuth: Authentication state management (token, role, tenant_id)
  • useDashboard: Dashboard data fetching and caching
  • useFindings: Findings list and detail fetching

Backend Components

API Routes

  • auth.py: Login, token refresh, user info
  • dashboard.py: Dashboard data aggregation
  • findings.py: Findings CRUD operations
  • reports.py: Audit report generation
  • attack_paths.py: Attack path visualization
  • footprint.py: Digital footprint data
  • ai.py: AI translation and coaching

Services

  • ai_translator.py: OpenAI/Anthropic integration for risk translation
  • risk_calculator.py: Risk score calculation algorithm
  • report_generator.py: PDF report generation with Jinja2/WeasyPrint

Core

  • config.py: Configuration management with Pydantic Settings
  • security.py: JWT token management, password hashing, RBAC decorators

Data Model

Entity Relationship Diagram

┌─────────────┐       ┌─────────────┐       ┌─────────────┐
│   tenants   │───────│    users    │───────│  findings   │
│─────────────│ 1:N   │─────────────│ 1:N   │─────────────│
│ id (PK)     │       │ id (PK)     │       │ id (PK)     │
│ name        │       │ tenant_id   │       │ tenant_id   │
│ slug        │       │ email       │       │ asset_id    │
│ industry    │       │ role        │       │ executive_id│
│ size_range  │       │ ...         │       │ severity    │
│ ...         │       └─────────────┘       │ status      │
└─────────────┘                            │ category    │
         │                                  │ ai_summary  │
         │                                  │ ...         │
         │                                  └─────────────┘
         │                                           │
         │                                           │
┌─────────────┐                            ┌───────────▼──────────┐
│   assets    │                            │    risk_scores       │
│─────────────│                            │──────────────────────│
│ id (PK)     │                            │ id (PK)             │
│ tenant_id   │                            │ tenant_id           │
│ name        │                            │ score_date          │
│ asset_type  │                            │ overall_score       │
│ value       │                            │ score_identity      │
│ ...         │                            │ score_cloud         │
└─────────────┘                            │ ...                 │
                                           └──────────────────────┘

┌─────────────┐       ┌─────────────┐       ┌─────────────┐
│ executives  │       │authorized   │       │attack_paths │
│─────────────│       │  assets    │       │─────────────│
│ id (PK)     │       │─────────────│       │ id (PK)     │
│ tenant_id   │       │ id (PK)     │       │ finding_id  │
│ full_name   │       │ tenant_id   │       │ title       │
│ title       │       │ value       │       │ ai_narrative│
│ email       │       │ asset_type  │       │ nodes_json  │
│ ...         │       │ ...         │       │ edges_json  │
└─────────────┘       └─────────────┘       └─────────────┘

┌─────────────┐
│audit_reports│
│─────────────│
│ id (PK)     │
│ tenant_id   │
│ title       │
│ report_date │
│ baseline_   │
│   score     │
│ pdf_path    │
│ ...         │
└─────────────┘

Key Tables

tenants

Organizational units with complete data isolation.

  • id: UUID primary key
  • name: Organization display name
  • slug: URL-friendly identifier (unique)
  • industry: Industry classification
  • size_range: Company size (SMB, mid-market, enterprise)
  • contact_email: Primary contact
  • is_active: Active status for soft deletes

users

User accounts with role-based access control.

  • id: UUID primary key
  • tenant_id: Foreign key to tenants
  • email: Unique email address
  • hashed_password: Bcrypt hash
  • role: enum (executive, it_admin, trustos_admin)
  • is_active: Account status

findings

Security vulnerabilities and exposures.

  • id: UUID primary key
  • tenant_id: Foreign key to tenants
  • asset_id: Foreign key to assets (nullable)
  • executive_id: Foreign key to executives (nullable)
  • title: Human-readable title
  • severity: enum (critical, high, medium, low, info)
  • status: enum (open, in_progress, resolved, verified)
  • category: enum (external_exposure, cloud_posture, credential_exposure, etc.)
  • technical_description: Raw technical details
  • cve_id: CVE identifier (if applicable)
  • cvss_score: CVSS score (if applicable)
  • ai_summary: AI-generated plain-English summary
  • ai_business_impact: AI-generated business impact
  • ai_remediation_steps: AI-generated fix steps
  • assignee_email: Assigned team member
  • due_date: Remediation deadline
  • is_top_risk: Flag for top 3 risks

risk_scores

Daily snapshots of risk metrics.

  • id: UUID primary key
  • tenant_id: Foreign key to tenants
  • score_date: Timestamp of snapshot
  • overall_score: 0-100 overall score
  • score_identity: Identity security score
  • score_cloud: Cloud posture score
  • score_network: Network security score
  • score_web: Web application score
  • score_credential: Credential security score
  • score_digital_footprint: Digital footprint score
  • score_third_party: Third-party risk score
  • critical_count: Count of critical findings
  • high_count: Count of high findings
  • medium_count: Count of medium findings
  • low_count: Count of low findings

Authentication & Authorization

Authentication Flow

1. User submits credentials to POST /api/v1/auth/login
   ↓
2. Backend validates credentials against database
   ↓
3. Backend generates JWT token with:
   - sub: user_id
   - role: user_role
   - tenant_id: tenant_id
   - exp: expiration timestamp
   ↓
4. Frontend stores token in localStorage
   ↓
5. Frontend includes token in Authorization header: Bearer <token>
   ↓
6. Backend validates token on each protected request
   ↓
7. Backend extracts user context from token
   ↓
8. Request proceeds with user context

JWT Token Structure

{
  "sub": "user-uuid",
  "role": "executive",
  "tenant_id": "tenant-uuid",
  "exp": 1234567890,
  "iat": 1234567890
}

Role-Based Access Control (RBAC)

Three roles with distinct permissions:

Executive

  • Can View: Dashboard, risk scores, AI summaries, trends
  • Cannot View: Raw CVE data, technical details, other tenants
  • Can Modify: None (read-only)

IT Admin

  • Can View: All findings, technical details, CVEs, remediation steps
  • Can Modify: Finding status, assignees, due dates, resolution notes
  • Cannot View: Other tenants' data

TrustOS Admin

  • Can View: All tenants, all data, system metrics
  • Can Modify: Tenant settings, authorized assets, audit reports
  • Can Manage: Users, roles, system configuration

Authorization Middleware

# Example from app/core/security.py
def require_executive_or_above(payload: dict = Depends(verify_token)):
    if payload.get("role") not in ["executive", "it_admin", "trustos_admin"]:
        raise HTTPException(status_code=403, detail="Insufficient permissions")
    return payload

def require_it_or_above(payload: dict = Depends(verify_token)):
    if payload.get("role") not in ["it_admin", "trustos_admin"]:
        raise HTTPException(status_code=403, detail="Insufficient permissions")
    return payload

def require_admin(payload: dict = Depends(verify_token)):
    if payload.get("role") != "trustos_admin":
        raise HTTPException(status_code=403, detail="Admin access required")
    return payload

API Design

RESTful Conventions

TrustOS follows RESTful API design principles:

  • Resource-Based URLs: /api/v1/findings, /api/v1/tenants
  • HTTP Methods: GET (read), POST (create), PATCH (update), DELETE (delete)
  • Status Codes: 200 (success), 201 (created), 400 (bad request), 401 (unauthorized), 403 (forbidden), 404 (not found), 500 (server error)
  • JSON Request/Response: All data is JSON-encoded
  • Versioning: /api/v1/ prefix for future compatibility

API Response Format

Success Response

{
  "data": { ... },
  "meta": {
    "timestamp": "2024-01-01T00:00:00Z",
    "request_id": "uuid"
  }
}

Error Response

{
  "error": {
    "code": "VALIDATION_ERROR",
    "message": "Invalid input data",
    "details": { ... }
  },
  "meta": {
    "timestamp": "2024-01-01T00:00:00Z",
    "request_id": "uuid"
  }
}

Key API Endpoints

Authentication

  • POST /api/v1/auth/login - Authenticate and receive token
  • GET /api/v1/auth/me - Get current user info

Dashboard

  • GET /api/v1/dashboard?tenant_id={id} - Get dashboard data

Findings

  • GET /api/v1/findings?tenant_id={id} - List findings
  • GET /api/v1/findings/{id} - Get finding detail
  • POST /api/v1/findings - Create finding (IT Admin+)
  • PATCH /api/v1/findings/{id}/status - Update finding status (IT Admin+)

Reports

  • GET /api/v1/audit-reports?tenant_id={id} - List audit reports
  • POST /api/v1/audit-reports/generate - Generate audit report (Admin)

AI

  • POST /api/v1/ai/translate/{finding_id} - Trigger AI translation (IT Admin+)
  • GET /api/v1/ai/explain/{finding_id}?question=... - AI Security Coach (IT Admin+)

Frontend Architecture

Next.js App Router Structure

src/app/
├── layout.tsx          # Root layout with providers
├── page.tsx            # Root redirect to /login
├── login/
│   └── page.tsx        # Login page
├── dashboard/
│   └── page.tsx        # Executive dashboard
├── findings/
│   ├── page.tsx        # Findings list
│   └── [id]/
│       └── page.tsx    # Finding detail
└── footprint/
    └── page.tsx        # Digital footprint center

State Management

TrustOS uses React Context for global state:

// Auth Context
interface AuthContext {
  token: string | null;
  role: UserRole | null;
  tenantId: string | null;
  name: string | null;
  login: (email: string, password: string) => Promise<void>;
  logout: () => void;
}

API Client

Custom API client with token management:

const api = {
  login: (email: string, password: string) => 
    fetch('/api/v1/auth/login', { method: 'POST', body: ... }),
  
  dashboard: (tenantId: string) => 
    fetch(`/api/v1/dashboard?tenant_id=${tenantId}`, {
      headers: { Authorization: `Bearer ${token}` }
    }),
  
  // ... other methods
};

Component Design Patterns

Presentational Components

  • Receive data via props
  • No side effects
  • Reusable across contexts

Container Components

  • Fetch data from API
  • Manage state
  • Pass data to presentational components

Higher-Order Components

  • withAuth: Wraps components requiring authentication
  • withRole: Wraps components requiring specific roles

Backend Architecture

FastAPI Application Structure

# app/main.py
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware

app = FastAPI(
    title="TrustOS API",
    version="1.0.0",
    description="AI-powered cyber resilience platform"
)

# CORS middleware
app.add_middleware(
    CORSMiddleware,
    allow_origins=["http://localhost:3000"],
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)

# Include routers
app.include_router(auth.router, prefix="/api/v1/auth", tags=["auth"])
app.include_router(dashboard.router, prefix="/api/v1/dashboard", tags=["dashboard"])
# ... other routers

@app.on_event("startup")
async def startup_event():
    # Create tables in development
    if settings.DEBUG:
        async with engine.begin() as conn:
            await conn.run_sync(Base.metadata.create_all)

@app.get("/health")
async def health_check():
    return {"status": "healthy"}

Dependency Injection

FastAPI's dependency system for authentication and database:

# Database session
async def get_db() -> AsyncSession:
    async with AsyncSessionLocal() as session:
        try:
            yield session
        finally:
            await session.close()

# Authentication
async def verify_token(authorization: str = Header(...)) -> dict:
    token = authorization.replace("Bearer ", "")
    payload = decode_jwt(token)
    return payload

# Protected route
@router.get("/dashboard")
async def get_dashboard(
    tenant_id: str,
    payload: dict = Depends(verify_token),
    db: AsyncSession = Depends(get_db)
):
    # Access payload['role'], payload['tenant_id']
    # Use db for database operations
    ...

Service Layer Pattern

Business logic separated from API routes:

# app/services/risk_calculator.py
async def recalculate_risk_score(tenant_id: str, db: AsyncSession) -> float:
    # Business logic for risk calculation
    findings = await get_open_findings(tenant_id, db)
    score = calculate_score(findings)
    await save_risk_snapshot(tenant_id, score, db)
    return score

# API route uses service
@router.patch("/findings/{id}/status")
async def update_status(finding_id: str, body: StatusUpdate, db: AsyncSession = Depends(get_db)):
    finding = await get_finding(finding_id, db)
    finding.status = body.status
    await db.commit()
    
    # Trigger background risk recalculation
    asyncio.create_task(recalculate_risk_score(finding.tenant_id))
    
    return finding

Database Design

Schema Design Principles

  1. Multi-Tenant by Default: All tables include tenant_id
  2. UUID Primary Keys: Distributed-friendly, no sequence contention
  3. Soft Deletes: is_active flags instead of hard deletes
  4. Audit Fields: created_at, updated_at on all tables
  5. JSONB for Flexibility: Store semi-structured data in JSONB columns

Indexing Strategy

-- Tenant isolation (every query)
CREATE INDEX idx_findings_tenant_id ON findings(tenant_id);

-- Common query patterns
CREATE INDEX idx_findings_status ON findings(status);
CREATE INDEX idx_findings_severity ON findings(severity);
CREATE INDEX idx_findings_category ON findings(category);
CREATE INDEX idx_findings_tenant_status ON findings(tenant_id, status);

-- Time-series queries
CREATE INDEX idx_risk_scores_tenant_date ON risk_scores(tenant_id, score_date DESC);

-- Unique constraints
CREATE UNIQUE INDEX idx_users_email ON users(email);
CREATE UNIQUE INDEX idx_tenants_slug ON tenants(slug);

Connection Pooling

# SQLAlchemy async engine with connection pooling
engine = create_async_engine(
    settings.DATABASE_URL,
    echo=False,
    pool_pre_ping=True,  # Verify connections before use
    pool_size=10,       # Base pool size
    max_overflow=20,     # Additional connections under load
)

AI Integration

AI Service Architecture

┌─────────────┐
│  API Route  │
└──────┬──────┘
       │
┌──────▼──────────┐
│ AI Translator   │
│ Service         │
└──────┬──────────┘
       │
┌──────▼──────────┐
│ LLM Provider    │
│ (OpenAI/Anthropic)│
└─────────────────┘

AI Translation Flow

# 1. Finding created or updated
finding = Finding(...)

# 2. Trigger AI translation (background task)
asyncio.create_task(translate_finding_async(finding.id))

# 3. AI service constructs prompt
prompt = f"""
Translate this cybersecurity finding:
Title: {finding.title}
Severity: {finding.severity}
CVE ID: {finding.cve_id}
Technical Description: {finding.technical_description}
"""

# 4. Call LLM with system prompt
system_prompt = """
You are TrustOS, an AI cyber resilience advisor.
Translate technical findings into plain-English business impact.
Output JSON with: summary, business_impact, impact_level, remediation_steps.
"""

# 5. Parse and store AI response
data = json.loads(llm_response)
finding.ai_summary = data["summary"]
finding.ai_business_impact = data["business_impact"]
finding.ai_remediation_steps = data["remediation_steps"]
await db.commit()

AI Provider Selection

# app/core/config.py
AI_PROVIDER = os.getenv("AI_PROVIDER", "openai")  # or "anthropic"

# app/services/ai_translator.py
if settings.AI_PROVIDER == "openai":
    client = AsyncOpenAI(api_key=settings.OPENAI_API_KEY)
    response = await client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[...],
        response_format={"type": "json_object"}
    )
elif settings.AI_PROVIDER == "anthropic":
    client = AsyncAnthropic(api_key=settings.ANTHROPIC_API_KEY)
    response = await client.messages.create(
        model="claude-3-haiku-20240307",
        messages=[...]
    )

AI Caching Strategy

  • Store AI Responses: AI translations are stored in the database
  • Avoid Redundant Calls: Check if ai_summary exists before re-translating
  • Background Processing: AI calls are async and non-blocking
  • Graceful Degradation: If AI fails, use raw technical description

Security Architecture

Defense in Depth

┌─────────────────────────────────────────────────────────┐
│ 1. Network Security                                      │
│    - HTTPS/TLS encryption                                │
│    - CORS restrictions                                   │
│    - Rate limiting (planned)                             │
└─────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ 2. Authentication                                         │
│    - JWT token validation                                 │
│    - Token expiration (8 hours)                          │
│    - Secure password hashing (bcrypt)                     │
└─────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ 3. Authorization                                          │
│    - Role-based access control                           │
│    - Tenant isolation enforcement                        │
│    - Route-level permission checks                       │
└─────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ 4. Input Validation                                       │
│    - Pydantic schema validation                          │
│    - SQL injection prevention (ORM)                      │
│    - XSS prevention (React escaping)                     │
└─────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ 5. Data Security                                         │
│    - Multi-tenant database isolation                     │
│    - Encrypted secrets management                        │
│    - Audit logging                                       │
└─────────────────────────────────────────────────────────┘

Security Headers

# app/main.py
from fastapi.middleware.trustedhost import TrustedHostMiddleware

app.add_middleware(
    TrustedHostMiddleware,
    allowed_hosts=["trustos.com", "*.trustos.com"]
)

# Additional headers via middleware
@app.middleware("http")
async def add_security_headers(request: Request, call_next):
    response = await call_next(request)
    response.headers["X-Content-Type-Options"] = "nosniff"
    response.headers["X-Frame-Options"] = "DENY"
    response.headers["X-XSS-Protection"] = "1; mode=block"
    response.headers["Strict-Transport-Security"] = "max-age=31536000; includeSubDomains"
    return response

Secrets Management

  • Environment Variables: All secrets in .env files (git-ignored)
  • Production: Use secret management services (AWS Secrets Manager, HashiCorp Vault)
  • Rotation: Regular API key rotation policy
  • Least Privilege: Database users have minimal required permissions

Scalability Considerations

Horizontal Scaling

The architecture supports horizontal scaling:

  • Stateless API: FastAPI instances can be scaled horizontally
  • Database Connection Pooling: Efficient connection reuse
  • Async I/O: High concurrency with minimal threads
  • Load Balancer: Nginx or cloud load balancer in front of API

Vertical Scaling

  • Database: PostgreSQL can scale vertically (more CPU/RAM)
  • Caching: Redis planned for session and query caching
  • CDN: Frontend static assets served via CDN

Performance Optimization Strategies

  1. Database Optimization

    • Proper indexing on frequently queried columns
    • Query optimization with EXPLAIN ANALYZE
    • Connection pooling to reduce overhead
    • Read replicas for read-heavy workloads (future)
  2. API Optimization

    • Response compression (gzip)
    • Pagination for large result sets
    • Selective field loading (avoid SELECT *)
    • Async operations throughout
  3. Frontend Optimization

    • Code splitting with Next.js
    • Image optimization
    • Static generation where possible
    • Client-side caching
  4. Caching Strategy

    • API response caching (Redis)
    • Static asset caching (CDN)
    • AI response caching (database)
    • Browser caching headers

Performance Optimization

Database Query Optimization

# Bad: N+1 query problem
for finding in findings:
    asset = await get_asset(finding.asset_id)  # N queries

# Good: Eager loading
findings = await db.execute(
    select(Finding).options(selectinload(Finding.asset))
)

API Response Optimization

# Bad: Return all fields
@router.get("/findings")
async def list_findings():
    return await db.execute(select(Finding))

# Good: Select only needed fields
@router.get("/findings")
async def list_findings():
    return await db.execute(
        select(Finding.id, Finding.title, Finding.severity)
    )

Frontend Performance

// Bad: Re-render on every state change
useEffect(() => {
  fetchData();
}, [state]); // Runs on every state change

// Good: Only re-fetch when dependencies change
useEffect(() => {
  fetchData();
}, [tenantId, filter]); // Only runs when tenant or filter changes

Monitoring & Observability

Logging Strategy

# Structured logging
import logging
logger = logging.getLogger(__name__)

logger.info(
    "Finding status updated",
    extra={
        "finding_id": finding.id,
        "old_status": old_status,
        "new_status": new_status,
        "user_id": user_id,
        "tenant_id": tenant_id
    }
)

Metrics to Track

  • API Metrics: Request rate, error rate, response time
  • Database Metrics: Query time, connection pool usage
  • Business Metrics: Active tenants, findings created, risk score trends
  • AI Metrics: API calls, token usage, cost tracking

Health Checks

@app.get("/health")
async def health_check():
    checks = {
        "database": await check_database(),
        "ai_service": await check_ai_service(),
        "storage": await check_storage()
    }
    status = "healthy" if all(checks.values()) else "degraded"
    return {"status": status, "checks": checks}

Future Architecture Enhancements

Planned Improvements

  1. Message Queue: Celery + Redis for background job processing
  2. Caching Layer: Redis for session and query caching
  3. Read Replicas: PostgreSQL read replicas for scaling reads
  4. Microservices: Split into separate services (auth, findings, monitoring)
  5. Event Sourcing: Event-driven architecture for audit trail
  6. GraphQL: Alternative to REST for complex queries
  7. Real-time Updates: WebSocket for live dashboard updates
  8. Edge Computing: Cloudflare Workers for global distribution

Conclusion

The TrustOS architecture is designed to be secure, scalable, and maintainable while following modern best practices. The multi-tenant isolation, AI augmentation, and authorization-first principles ensure the platform can grow with customer needs while maintaining security and privacy.

For questions or contributions to the architecture, please refer to the CONTRIBUTING.md guide.