# TrustOS Architecture Documentation This document provides a detailed overview of the TrustOS system architecture, design decisions, data flow, and technical implementation details. --- ## Table of Contents - [System Overview](#system-overview) - [Architecture Principles](#architecture-principles) - [Component Architecture](#component-architecture) - [Data Model](#data-model) - [Authentication & Authorization](#authentication--authorization) - [API Design](#api-design) - [Frontend Architecture](#frontend-architecture) - [Backend Architecture](#backend-architecture) - [Database Design](#database-design) - [AI Integration](#ai-integration) - [Security Architecture](#security-architecture) - [Scalability Considerations](#scalability-considerations) - [Performance Optimization](#performance-optimization) --- ## System Overview TrustOS is a multi-tenant, AI-powered cyber resilience platform built on a modern microservices-inspired architecture. The system is designed to be: - **Secure**: Multi-tenant isolation with defense-in-depth security - **Scalable**: Async I/O throughout for high concurrency - **Maintainable**: Clean separation of concerns and modular design - **Resilient**: Graceful degradation when external services are unavailable ### High-Level Architecture ``` ┌─────────────────────────────────────────────────────────────────┐ │ Client Layer │ │ Web Browser (Executive, IT Admin, TrustOS Admin) │ └────────────────────┬────────────────────────────────────────────┘ │ HTTPS ┌────────────────────▼────────────────────────────────────────────┐ │ Frontend Layer │ │ Next.js 16 + TypeScript + Tailwind CSS + shadcn/ui │ │ - Server-Side Rendering (SSR) │ │ - Client-Side Hydration │ │ - Static Site Generation (SSG) where applicable │ └────────────────────┬────────────────────────────────────────────┘ │ REST API (JSON) ┌────────────────────▼────────────────────────────────────────────┐ │ API Gateway │ │ FastAPI Application │ │ - Request Validation (Pydantic) │ │ - Authentication (JWT) │ │ - Authorization (RBAC) │ │ - Rate Limiting (future) │ │ - Request Logging │ └────────────────────┬────────────────────────────────────────────┘ │ ┌────────────┴────────────┐ │ │ ┌───────▼────────┐ ┌────────▼─────────┐ │ Service Layer │ │ Background │ │ │ │ Workers │ │ - Dashboard │ │ - AI Translation│ │ - Findings │ │ - Risk Calc │ │ - Reports │ │ - PDF Gen │ │ - Footprint │ │ - Monitoring │ └───────┬────────┘ └────────┬─────────┘ │ │ ┌───────▼─────────────────────────▼──────────┐ │ Data Access Layer │ │ SQLAlchemy 2.0 (Async ORM) │ │ - Query Building │ │ - Connection Pooling │ │ - Transaction Management │ └────────────────────┬─────────────────────────┘ │ ┌────────────────────▼─────────────────────────┐ │ Database Layer │ │ PostgreSQL 16 │ │ - Multi-Tenant Data Isolation │ │ - Indexing Strategy │ │ - Foreign Key Constraints │ │ - JSONB for Flexible Data │ └──────────────────────────────────────────────┘ External Services: ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ OpenAI API │ │ Anthropic API│ │ HIBP API │ └──────────────┘ └──────────────┘ └──────────────┘ ``` --- ## Architecture Principles ### 1. Authorization First TrustOS only scans and monitors assets that have been explicitly authorized by the tenant organization. This is a core security and privacy principle: - **Authorized Assets Table**: All monitored assets must be pre-registered - **Scope Enforcement**: All automated checks respect the authorized asset list - **Executive Enrollment**: Executive monitoring requires explicit organizational consent - **Audit Trail**: All authorization decisions are logged ### 2. Multi-Tenant Isolation Every data access path enforces tenant isolation at multiple layers: - **Database Level**: All tables include `tenant_id` with foreign key constraints - **ORM Level**: Queries automatically filter by tenant_id - **API Level**: Middleware validates tenant access before processing requests - **Application Level**: UI components only display tenant-specific data ### 3. Privacy by Design Executive and organizational data is handled with privacy as a foundational requirement: - **Minimal Data Collection**: Only collect data necessary for security assessments - **Explicit Consent**: Executive enrollment requires organizational authorization - **Data Minimization**: Store only what, not who, where possible - **Audit Logging**: All data access is logged for accountability ### 4. AI-Augmented, Not AI-Dependent AI features enhance the product but are not required for core functionality: - **Graceful Degradation**: Features work with raw technical data if AI is unavailable - **Caching**: AI responses are cached to avoid redundant API calls - **Fallback Mechanisms**: System continues operating if AI services are down - **Cost Control**: Rate limiting and caching to manage AI API costs ### 5. Audit Trail All state changes are tracked with full provenance: - **Timestamps**: `created_at` and `updated_at` on all records - **User Attribution**: `created_by` and `updated_by` where applicable - **Status Changes**: Finding status transitions are logged with notes - **Access Logs**: API requests are logged with user and tenant context --- ## Component Architecture ### Frontend Components #### Page Components - **Dashboard Page** (`/dashboard`): Executive view with risk score, top risks, trends - **Findings Page** (`/findings`): IT admin view with sortable/filterable table - **Finding Detail Page** (`/findings/[id]`): Detailed view with technical and business impact - **Login Page** (`/login`): Authentication interface - **Footprint Page** (`/footprint`): Digital footprint center #### Reusable Components) - **RiskDial**: Circular gauge for cyber health score (0-100) - **RiskCard**: Card displaying finding with AI summary and impact - **TrendChart**: Line chart for 90-day risk history - **RemediationBoard**: Kanban board for finding status tracking - **Badge**: Severity and status badges with color coding #### Hooks - **useAuth**: Authentication state management (token, role, tenant_id) - **useDashboard**: Dashboard data fetching and caching - **useFindings**: Findings list and detail fetching ### Backend Components #### API Routes - **auth.py**: Login, token refresh, user info - **dashboard.py**: Dashboard data aggregation - **findings.py**: Findings CRUD operations - **reports.py**: Audit report generation - **attack_paths.py**: Attack path visualization - **footprint.py**: Digital footprint data - **ai.py**: AI translation and coaching #### Services - **ai_translator.py**: OpenAI/Anthropic integration for risk translation - **risk_calculator.py**: Risk score calculation algorithm - **report_generator.py**: PDF report generation with Jinja2/WeasyPrint #### Core - **config.py**: Configuration management with Pydantic Settings - **security.py**: JWT token management, password hashing, RBAC decorators --- ## Data Model ### Entity Relationship Diagram ``` ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ tenants │───────│ users │───────│ findings │ │─────────────│ 1:N │─────────────│ 1:N │─────────────│ │ id (PK) │ │ id (PK) │ │ id (PK) │ │ name │ │ tenant_id │ │ tenant_id │ │ slug │ │ email │ │ asset_id │ │ industry │ │ role │ │ executive_id│ │ size_range │ │ ... │ │ severity │ │ ... │ └─────────────┘ │ status │ └─────────────┘ │ category │ │ │ ai_summary │ │ │ ... │ │ └─────────────┘ │ │ │ │ ┌─────────────┐ ┌───────────▼──────────┐ │ assets │ │ risk_scores │ │─────────────│ │──────────────────────│ │ id (PK) │ │ id (PK) │ │ tenant_id │ │ tenant_id │ │ name │ │ score_date │ │ asset_type │ │ overall_score │ │ value │ │ score_identity │ │ ... │ │ score_cloud │ └─────────────┘ │ ... │ └──────────────────────┘ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ executives │ │authorized │ │attack_paths │ │─────────────│ │ assets │ │─────────────│ │ id (PK) │ │─────────────│ │ id (PK) │ │ tenant_id │ │ id (PK) │ │ finding_id │ │ full_name │ │ tenant_id │ │ title │ │ title │ │ value │ │ ai_narrative│ │ email │ │ asset_type │ │ nodes_json │ │ ... │ │ ... │ │ edges_json │ └─────────────┘ └─────────────┘ └─────────────┘ ┌─────────────┐ │audit_reports│ │─────────────│ │ id (PK) │ │ tenant_id │ │ title │ │ report_date │ │ baseline_ │ │ score │ │ pdf_path │ │ ... │ └─────────────┘ ``` ### Key Tables #### tenants Organizational units with complete data isolation. - `id`: UUID primary key - `name`: Organization display name - `slug`: URL-friendly identifier (unique) - `industry`: Industry classification - `size_range`: Company size (SMB, mid-market, enterprise) - `contact_email`: Primary contact - `is_active`: Active status for soft deletes #### users User accounts with role-based access control. - `id`: UUID primary key - `tenant_id`: Foreign key to tenants - `email`: Unique email address - `hashed_password`: Bcrypt hash - `role`: enum (executive, it_admin, trustos_admin) - `is_active`: Account status #### findings Security vulnerabilities and exposures. - `id`: UUID primary key - `tenant_id`: Foreign key to tenants - `asset_id`: Foreign key to assets (nullable) - `executive_id`: Foreign key to executives (nullable) - `title`: Human-readable title - `severity`: enum (critical, high, medium, low, info) - `status`: enum (open, in_progress, resolved, verified) - `category`: enum (external_exposure, cloud_posture, credential_exposure, etc.) - `technical_description`: Raw technical details - `cve_id`: CVE identifier (if applicable) - `cvss_score`: CVSS score (if applicable) - `ai_summary`: AI-generated plain-English summary - `ai_business_impact`: AI-generated business impact - `ai_remediation_steps`: AI-generated fix steps - `assignee_email`: Assigned team member - `due_date`: Remediation deadline - `is_top_risk`: Flag for top 3 risks #### risk_scores Daily snapshots of risk metrics. - `id`: UUID primary key - `tenant_id`: Foreign key to tenants - `score_date`: Timestamp of snapshot - `overall_score`: 0-100 overall score - `score_identity`: Identity security score - `score_cloud`: Cloud posture score - `score_network`: Network security score - `score_web`: Web application score - `score_credential`: Credential security score - `score_digital_footprint`: Digital footprint score - `score_third_party`: Third-party risk score - `critical_count`: Count of critical findings - `high_count`: Count of high findings - `medium_count`: Count of medium findings - `low_count`: Count of low findings --- ## Authentication & Authorization ### Authentication Flow ``` 1. User submits credentials to POST /api/v1/auth/login ↓ 2. Backend validates credentials against database ↓ 3. Backend generates JWT token with: - sub: user_id - role: user_role - tenant_id: tenant_id - exp: expiration timestamp ↓ 4. Frontend stores token in localStorage ↓ 5. Frontend includes token in Authorization header: Bearer ↓ 6. Backend validates token on each protected request ↓ 7. Backend extracts user context from token ↓ 8. Request proceeds with user context ``` ### JWT Token Structure ```json { "sub": "user-uuid", "role": "executive", "tenant_id": "tenant-uuid", "exp": 1234567890, "iat": 1234567890 } ``` ### Role-Based Access Control (RBAC) Three roles with distinct permissions: #### Executive - **Can View**: Dashboard, risk scores, AI summaries, trends - **Cannot View**: Raw CVE data, technical details, other tenants - **Can Modify**: None (read-only) #### IT Admin - **Can View**: All findings, technical details, CVEs, remediation steps - **Can Modify**: Finding status, assignees, due dates, resolution notes - **Cannot View**: Other tenants' data #### TrustOS Admin - **Can View**: All tenants, all data, system metrics - **Can Modify**: Tenant settings, authorized assets, audit reports - **Can Manage**: Users, roles, system configuration ### Authorization Middleware ```python # Example from app/core/security.py def require_executive_or_above(payload: dict = Depends(verify_token)): if payload.get("role") not in ["executive", "it_admin", "trustos_admin"]: raise HTTPException(status_code=403, detail="Insufficient permissions") return payload def require_it_or_above(payload: dict = Depends(verify_token)): if payload.get("role") not in ["it_admin", "trustos_admin"]: raise HTTPException(status_code=403, detail="Insufficient permissions") return payload def require_admin(payload: dict = Depends(verify_token)): if payload.get("role") != "trustos_admin": raise HTTPException(status_code=403, detail="Admin access required") return payload ``` --- ## API Design ### RESTful Conventions TrustOS follows RESTful API design principles: - **Resource-Based URLs**: `/api/v1/findings`, `/api/v1/tenants` - **HTTP Methods**: GET (read), POST (create), PATCH (update), DELETE (delete) - **Status Codes**: 200 (success), 201 (created), 400 (bad request), 401 (unauthorized), 403 (forbidden), 404 (not found), 500 (server error) - **JSON Request/Response**: All data is JSON-encoded - **Versioning**: `/api/v1/` prefix for future compatibility ### API Response Format #### Success Response ```json { "data": { ... }, "meta": { "timestamp": "2024-01-01T00:00:00Z", "request_id": "uuid" } } ``` #### Error Response ```json { "error": { "code": "VALIDATION_ERROR", "message": "Invalid input data", "details": { ... } }, "meta": { "timestamp": "2024-01-01T00:00:00Z", "request_id": "uuid" } } ``` ### Key API Endpoints #### Authentication - `POST /api/v1/auth/login` - Authenticate and receive token - `GET /api/v1/auth/me` - Get current user info #### Dashboard - `GET /api/v1/dashboard?tenant_id={id}` - Get dashboard data #### Findings - `GET /api/v1/findings?tenant_id={id}` - List findings - `GET /api/v1/findings/{id}` - Get finding detail - `POST /api/v1/findings` - Create finding (IT Admin+) - `PATCH /api/v1/findings/{id}/status` - Update finding status (IT Admin+) #### Reports - `GET /api/v1/audit-reports?tenant_id={id}` - List audit reports - `POST /api/v1/audit-reports/generate` - Generate audit report (Admin) #### AI - `POST /api/v1/ai/translate/{finding_id}` - Trigger AI translation (IT Admin+) - `GET /api/v1/ai/explain/{finding_id}?question=...` - AI Security Coach (IT Admin+) --- ## Frontend Architecture ### Next.js App Router Structure ``` src/app/ ├── layout.tsx # Root layout with providers ├── page.tsx # Root redirect to /login ├── login/ │ └── page.tsx # Login page ├── dashboard/ │ └── page.tsx # Executive dashboard ├── findings/ │ ├── page.tsx # Findings list │ └── [id]/ │ └── page.tsx # Finding detail └── footprint/ └── page.tsx # Digital footprint center ``` ### State Management TrustOS uses React Context for global state: ```typescript // Auth Context interface AuthContext { token: string | null; role: UserRole | null; tenantId: string | null; name: string | null; login: (email: string, password: string) => Promise; logout: () => void; } ``` ### API Client Custom API client with token management: ```typescript const api = { login: (email: string, password: string) => fetch('/api/v1/auth/login', { method: 'POST', body: ... }), dashboard: (tenantId: string) => fetch(`/api/v1/dashboard?tenant_id=${tenantId}`, { headers: { Authorization: `Bearer ${token}` } }), // ... other methods }; ``` ### Component Design Patterns #### Presentational Components - Receive data via props - No side effects - Reusable across contexts #### Container Components - Fetch data from API - Manage state - Pass data to presentational components #### Higher-Order Components - `withAuth`: Wraps components requiring authentication - `withRole`: Wraps components requiring specific roles --- ## Backend Architecture ### FastAPI Application Structure ```python # app/main.py from fastapi import FastAPI from fastapi.middleware.cors import CORSMiddleware app = FastAPI( title="TrustOS API", version="1.0.0", description="AI-powered cyber resilience platform" ) # CORS middleware app.add_middleware( CORSMiddleware, allow_origins=["http://localhost:3000"], allow_credentials=True, allow_methods=["*"], allow_headers=["*"], ) # Include routers app.include_router(auth.router, prefix="/api/v1/auth", tags=["auth"]) app.include_router(dashboard.router, prefix="/api/v1/dashboard", tags=["dashboard"]) # ... other routers @app.on_event("startup") async def startup_event(): # Create tables in development if settings.DEBUG: async with engine.begin() as conn: await conn.run_sync(Base.metadata.create_all) @app.get("/health") async def health_check(): return {"status": "healthy"} ``` ### Dependency Injection FastAPI's dependency system for authentication and database: ```python # Database session async def get_db() -> AsyncSession: async with AsyncSessionLocal() as session: try: yield session finally: await session.close() # Authentication async def verify_token(authorization: str = Header(...)) -> dict: token = authorization.replace("Bearer ", "") payload = decode_jwt(token) return payload # Protected route @router.get("/dashboard") async def get_dashboard( tenant_id: str, payload: dict = Depends(verify_token), db: AsyncSession = Depends(get_db) ): # Access payload['role'], payload['tenant_id'] # Use db for database operations ... ``` ### Service Layer Pattern Business logic separated from API routes: ```python # app/services/risk_calculator.py async def recalculate_risk_score(tenant_id: str, db: AsyncSession) -> float: # Business logic for risk calculation findings = await get_open_findings(tenant_id, db) score = calculate_score(findings) await save_risk_snapshot(tenant_id, score, db) return score # API route uses service @router.patch("/findings/{id}/status") async def update_status(finding_id: str, body: StatusUpdate, db: AsyncSession = Depends(get_db)): finding = await get_finding(finding_id, db) finding.status = body.status await db.commit() # Trigger background risk recalculation asyncio.create_task(recalculate_risk_score(finding.tenant_id)) return finding ``` --- ## Database Design ### Schema Design Principles 1. **Multi-Tenant by Default**: All tables include `tenant_id` 2. **UUID Primary Keys**: Distributed-friendly, no sequence contention 3. **Soft Deletes**: `is_active` flags instead of hard deletes 4. **Audit Fields**: `created_at`, `updated_at` on all tables 5. **JSONB for Flexibility**: Store semi-structured data in JSONB columns ### Indexing Strategy ```sql -- Tenant isolation (every query) CREATE INDEX idx_findings_tenant_id ON findings(tenant_id); -- Common query patterns CREATE INDEX idx_findings_status ON findings(status); CREATE INDEX idx_findings_severity ON findings(severity); CREATE INDEX idx_findings_category ON findings(category); CREATE INDEX idx_findings_tenant_status ON findings(tenant_id, status); -- Time-series queries CREATE INDEX idx_risk_scores_tenant_date ON risk_scores(tenant_id, score_date DESC); -- Unique constraints CREATE UNIQUE INDEX idx_users_email ON users(email); CREATE UNIQUE INDEX idx_tenants_slug ON tenants(slug); ``` ### Connection Pooling ```python # SQLAlchemy async engine with connection pooling engine = create_async_engine( settings.DATABASE_URL, echo=False, pool_pre_ping=True, # Verify connections before use pool_size=10, # Base pool size max_overflow=20, # Additional connections under load ) ``` --- ## AI Integration ### AI Service Architecture ``` ┌─────────────┐ │ API Route │ └──────┬──────┘ │ ┌──────▼──────────┐ │ AI Translator │ │ Service │ └──────┬──────────┘ │ ┌──────▼──────────┐ │ LLM Provider │ │ (OpenAI/Anthropic)│ └─────────────────┘ ``` ### AI Translation Flow ```python # 1. Finding created or updated finding = Finding(...) # 2. Trigger AI translation (background task) asyncio.create_task(translate_finding_async(finding.id)) # 3. AI service constructs prompt prompt = f""" Translate this cybersecurity finding: Title: {finding.title} Severity: {finding.severity} CVE ID: {finding.cve_id} Technical Description: {finding.technical_description} """ # 4. Call LLM with system prompt system_prompt = """ You are TrustOS, an AI cyber resilience advisor. Translate technical findings into plain-English business impact. Output JSON with: summary, business_impact, impact_level, remediation_steps. """ # 5. Parse and store AI response data = json.loads(llm_response) finding.ai_summary = data["summary"] finding.ai_business_impact = data["business_impact"] finding.ai_remediation_steps = data["remediation_steps"] await db.commit() ``` ### AI Provider Selection ```python # app/core/config.py AI_PROVIDER = os.getenv("AI_PROVIDER", "openai") # or "anthropic" # app/services/ai_translator.py if settings.AI_PROVIDER == "openai": client = AsyncOpenAI(api_key=settings.OPENAI_API_KEY) response = await client.chat.completions.create( model="gpt-4o-mini", messages=[...], response_format={"type": "json_object"} ) elif settings.AI_PROVIDER == "anthropic": client = AsyncAnthropic(api_key=settings.ANTHROPIC_API_KEY) response = await client.messages.create( model="claude-3-haiku-20240307", messages=[...] ) ``` ### AI Caching Strategy - **Store AI Responses**: AI translations are stored in the database - **Avoid Redundant Calls**: Check if `ai_summary` exists before re-translating - **Background Processing**: AI calls are async and non-blocking - **Graceful Degradation**: If AI fails, use raw technical description --- ## Security Architecture ### Defense in Depth ``` ┌─────────────────────────────────────────────────────────┐ │ 1. Network Security │ │ - HTTPS/TLS encryption │ │ - CORS restrictions │ │ - Rate limiting (planned) │ └─────────────────────────────────────────────────────────┘ ┌─────────────────────────────────────────────────────────┐ │ 2. Authentication │ │ - JWT token validation │ │ - Token expiration (8 hours) │ │ - Secure password hashing (bcrypt) │ └─────────────────────────────────────────────────────────┘ ┌─────────────────────────────────────────────────────────┐ │ 3. Authorization │ │ - Role-based access control │ │ - Tenant isolation enforcement │ │ - Route-level permission checks │ └─────────────────────────────────────────────────────────┘ ┌─────────────────────────────────────────────────────────┐ │ 4. Input Validation │ │ - Pydantic schema validation │ │ - SQL injection prevention (ORM) │ │ - XSS prevention (React escaping) │ └─────────────────────────────────────────────────────────┘ ┌─────────────────────────────────────────────────────────┐ │ 5. Data Security │ │ - Multi-tenant database isolation │ │ - Encrypted secrets management │ │ - Audit logging │ └─────────────────────────────────────────────────────────┘ ``` ### Security Headers ```python # app/main.py from fastapi.middleware.trustedhost import TrustedHostMiddleware app.add_middleware( TrustedHostMiddleware, allowed_hosts=["trustos.com", "*.trustos.com"] ) # Additional headers via middleware @app.middleware("http") async def add_security_headers(request: Request, call_next): response = await call_next(request) response.headers["X-Content-Type-Options"] = "nosniff" response.headers["X-Frame-Options"] = "DENY" response.headers["X-XSS-Protection"] = "1; mode=block" response.headers["Strict-Transport-Security"] = "max-age=31536000; includeSubDomains" return response ``` ### Secrets Management - **Environment Variables**: All secrets in `.env` files (git-ignored) - **Production**: Use secret management services (AWS Secrets Manager, HashiCorp Vault) - **Rotation**: Regular API key rotation policy - **Least Privilege**: Database users have minimal required permissions --- ## Scalability Considerations ### Horizontal Scaling The architecture supports horizontal scaling: - **Stateless API**: FastAPI instances can be scaled horizontally - **Database Connection Pooling**: Efficient connection reuse - **Async I/O**: High concurrency with minimal threads - **Load Balancer**: Nginx or cloud load balancer in front of API ### Vertical Scaling - **Database**: PostgreSQL can scale vertically (more CPU/RAM) - **Caching**: Redis planned for session and query caching - **CDN**: Frontend static assets served via CDN ### Performance Optimization Strategies 1. **Database Optimization** - Proper indexing on frequently queried columns - Query optimization with EXPLAIN ANALYZE - Connection pooling to reduce overhead - Read replicas for read-heavy workloads (future) 2. **API Optimization** - Response compression (gzip) - Pagination for large result sets - Selective field loading (avoid SELECT *) - Async operations throughout 3. **Frontend Optimization** - Code splitting with Next.js - Image optimization - Static generation where possible - Client-side caching 4. **Caching Strategy** - API response caching (Redis) - Static asset caching (CDN) - AI response caching (database) - Browser caching headers --- ## Performance Optimization ### Database Query Optimization ```python # Bad: N+1 query problem for finding in findings: asset = await get_asset(finding.asset_id) # N queries # Good: Eager loading findings = await db.execute( select(Finding).options(selectinload(Finding.asset)) ) ``` ### API Response Optimization ```python # Bad: Return all fields @router.get("/findings") async def list_findings(): return await db.execute(select(Finding)) # Good: Select only needed fields @router.get("/findings") async def list_findings(): return await db.execute( select(Finding.id, Finding.title, Finding.severity) ) ``` ### Frontend Performance ```typescript // Bad: Re-render on every state change useEffect(() => { fetchData(); }, [state]); // Runs on every state change // Good: Only re-fetch when dependencies change useEffect(() => { fetchData(); }, [tenantId, filter]); // Only runs when tenant or filter changes ``` --- ## Monitoring & Observability ### Logging Strategy ```python # Structured logging import logging logger = logging.getLogger(__name__) logger.info( "Finding status updated", extra={ "finding_id": finding.id, "old_status": old_status, "new_status": new_status, "user_id": user_id, "tenant_id": tenant_id } ) ``` ### Metrics to Track - **API Metrics**: Request rate, error rate, response time - **Database Metrics**: Query time, connection pool usage - **Business Metrics**: Active tenants, findings created, risk score trends - **AI Metrics**: API calls, token usage, cost tracking ### Health Checks ```python @app.get("/health") async def health_check(): checks = { "database": await check_database(), "ai_service": await check_ai_service(), "storage": await check_storage() } status = "healthy" if all(checks.values()) else "degraded" return {"status": status, "checks": checks} ``` --- ## Future Architecture Enhancements ### Planned Improvements 1. **Message Queue**: Celery + Redis for background job processing 2. **Caching Layer**: Redis for session and query caching 3. **Read Replicas**: PostgreSQL read replicas for scaling reads 4. **Microservices**: Split into separate services (auth, findings, monitoring) 5. **Event Sourcing**: Event-driven architecture for audit trail 6. **GraphQL**: Alternative to REST for complex queries 7. **Real-time Updates**: WebSocket for live dashboard updates 8. **Edge Computing**: Cloudflare Workers for global distribution --- ## Conclusion The TrustOS architecture is designed to be secure, scalable, and maintainable while following modern best practices. The multi-tenant isolation, AI augmentation, and authorization-first principles ensure the platform can grow with customer needs while maintaining security and privacy. For questions or contributions to the architecture, please refer to the [CONTRIBUTING.md](CONTRIBUTING.md) guide.