docs: Add comprehensive documentation suite
- Enhance README.md with full installation instructions, architecture overview, features, configuration, testing, deployment, security, and troubleshooting sections - Add ARCHITECTURE.md with detailed system architecture, data model, authentication, API design, frontend/backend architecture, database design, AI integration, security, and scalability considerations -Add API.md with complete API reference including authentication, all endpoints, data models, examples, and interactive documentation links - Add DEPLOYMENT.md with deployment guides for Railway, Render, VPS, and Kubernetes, including pre-deployment checklist, monitoring, backup, and troubleshooting - Add CONTRIBUTING.md with development workflow, coding standards, testing guidelines, documentation standards, PR process, and community guidelines
This commit is contained in:
980
docs/ARCHITECTURE.md
Normal file
980
docs/ARCHITECTURE.md
Normal file
@@ -0,0 +1,980 @@
|
||||
# TrustOS Architecture Documentation
|
||||
|
||||
This document provides a detailed overview of the TrustOS system architecture, design decisions, data flow, and technical implementation details.
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
- [System Overview](#system-overview)
|
||||
- [Architecture Principles](#architecture-principles)
|
||||
- [Component Architecture](#component-architecture)
|
||||
- [Data Model](#data-model)
|
||||
- [Authentication & Authorization](#authentication--authorization)
|
||||
- [API Design](#api-design)
|
||||
- [Frontend Architecture](#frontend-architecture)
|
||||
- [Backend Architecture](#backend-architecture)
|
||||
- [Database Design](#database-design)
|
||||
- [AI Integration](#ai-integration)
|
||||
- [Security Architecture](#security-architecture)
|
||||
- [Scalability Considerations](#scalability-considerations)
|
||||
- [Performance Optimization](#performance-optimization)
|
||||
|
||||
---
|
||||
|
||||
## System Overview
|
||||
|
||||
TrustOS is a multi-tenant, AI-powered cyber resilience platform built on a modern microservices-inspired architecture. The system is designed to be:
|
||||
|
||||
- **Secure**: Multi-tenant isolation with defense-in-depth security
|
||||
- **Scalable**: Async I/O throughout for high concurrency
|
||||
- **Maintainable**: Clean separation of concerns and modular design
|
||||
- **Resilient**: Graceful degradation when external services are unavailable
|
||||
|
||||
### High-Level Architecture
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────┐
|
||||
│ Client Layer │
|
||||
│ Web Browser (Executive, IT Admin, TrustOS Admin) │
|
||||
└────────────────────┬────────────────────────────────────────────┘
|
||||
│ HTTPS
|
||||
┌────────────────────▼────────────────────────────────────────────┐
|
||||
│ Frontend Layer │
|
||||
│ Next.js 16 + TypeScript + Tailwind CSS + shadcn/ui │
|
||||
│ - Server-Side Rendering (SSR) │
|
||||
│ - Client-Side Hydration │
|
||||
│ - Static Site Generation (SSG) where applicable │
|
||||
└────────────────────┬────────────────────────────────────────────┘
|
||||
│ REST API (JSON)
|
||||
┌────────────────────▼────────────────────────────────────────────┐
|
||||
│ API Gateway │
|
||||
│ FastAPI Application │
|
||||
│ - Request Validation (Pydantic) │
|
||||
│ - Authentication (JWT) │
|
||||
│ - Authorization (RBAC) │
|
||||
│ - Rate Limiting (future) │
|
||||
│ - Request Logging │
|
||||
└────────────────────┬────────────────────────────────────────────┘
|
||||
│
|
||||
┌────────────┴────────────┐
|
||||
│ │
|
||||
┌───────▼────────┐ ┌────────▼─────────┐
|
||||
│ Service Layer │ │ Background │
|
||||
│ │ │ Workers │
|
||||
│ - Dashboard │ │ - AI Translation│
|
||||
│ - Findings │ │ - Risk Calc │
|
||||
│ - Reports │ │ - PDF Gen │
|
||||
│ - Footprint │ │ - Monitoring │
|
||||
└───────┬────────┘ └────────┬─────────┘
|
||||
│ │
|
||||
┌───────▼─────────────────────────▼──────────┐
|
||||
│ Data Access Layer │
|
||||
│ SQLAlchemy 2.0 (Async ORM) │
|
||||
│ - Query Building │
|
||||
│ - Connection Pooling │
|
||||
│ - Transaction Management │
|
||||
└────────────────────┬─────────────────────────┘
|
||||
│
|
||||
┌────────────────────▼─────────────────────────┐
|
||||
│ Database Layer │
|
||||
│ PostgreSQL 16 │
|
||||
│ - Multi-Tenant Data Isolation │
|
||||
│ - Indexing Strategy │
|
||||
│ - Foreign Key Constraints │
|
||||
│ - JSONB for Flexible Data │
|
||||
└──────────────────────────────────────────────┘
|
||||
|
||||
External Services:
|
||||
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
|
||||
│ OpenAI API │ │ Anthropic API│ │ HIBP API │
|
||||
└──────────────┘ └──────────────┘ └──────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Architecture Principles
|
||||
|
||||
### 1. Authorization First
|
||||
|
||||
TrustOS only scans and monitors assets that have been explicitly authorized by the tenant organization. This is a core security and privacy principle:
|
||||
|
||||
- **Authorized Assets Table**: All monitored assets must be pre-registered
|
||||
- **Scope Enforcement**: All automated checks respect the authorized asset list
|
||||
- **Executive Enrollment**: Executive monitoring requires explicit organizational consent
|
||||
- **Audit Trail**: All authorization decisions are logged
|
||||
|
||||
### 2. Multi-Tenant Isolation
|
||||
|
||||
Every data access path enforces tenant isolation at multiple layers:
|
||||
|
||||
- **Database Level**: All tables include `tenant_id` with foreign key constraints
|
||||
- **ORM Level**: Queries automatically filter by tenant_id
|
||||
- **API Level**: Middleware validates tenant access before processing requests
|
||||
- **Application Level**: UI components only display tenant-specific data
|
||||
|
||||
### 3. Privacy by Design
|
||||
|
||||
Executive and organizational data is handled with privacy as a foundational requirement:
|
||||
|
||||
- **Minimal Data Collection**: Only collect data necessary for security assessments
|
||||
- **Explicit Consent**: Executive enrollment requires organizational authorization
|
||||
- **Data Minimization**: Store only what, not who, where possible
|
||||
- **Audit Logging**: All data access is logged for accountability
|
||||
|
||||
### 4. AI-Augmented, Not AI-Dependent
|
||||
|
||||
AI features enhance the product but are not required for core functionality:
|
||||
|
||||
- **Graceful Degradation**: Features work with raw technical data if AI is unavailable
|
||||
- **Caching**: AI responses are cached to avoid redundant API calls
|
||||
- **Fallback Mechanisms**: System continues operating if AI services are down
|
||||
- **Cost Control**: Rate limiting and caching to manage AI API costs
|
||||
|
||||
### 5. Audit Trail
|
||||
|
||||
All state changes are tracked with full provenance:
|
||||
|
||||
- **Timestamps**: `created_at` and `updated_at` on all records
|
||||
- **User Attribution**: `created_by` and `updated_by` where applicable
|
||||
- **Status Changes**: Finding status transitions are logged with notes
|
||||
- **Access Logs**: API requests are logged with user and tenant context
|
||||
|
||||
---
|
||||
|
||||
## Component Architecture
|
||||
|
||||
### Frontend Components
|
||||
|
||||
#### Page Components
|
||||
- **Dashboard Page** (`/dashboard`): Executive view with risk score, top risks, trends
|
||||
- **Findings Page** (`/findings`): IT admin view with sortable/filterable table
|
||||
- **Finding Detail Page** (`/findings/[id]`): Detailed view with technical and business impact
|
||||
- **Login Page** (`/login`): Authentication interface
|
||||
- **Footprint Page** (`/footprint`): Digital footprint center
|
||||
|
||||
#### Reusable Components)
|
||||
- **RiskDial**: Circular gauge for cyber health score (0-100)
|
||||
- **RiskCard**: Card displaying finding with AI summary and impact
|
||||
- **TrendChart**: Line chart for 90-day risk history
|
||||
- **RemediationBoard**: Kanban board for finding status tracking
|
||||
- **Badge**: Severity and status badges with color coding
|
||||
|
||||
#### Hooks
|
||||
- **useAuth**: Authentication state management (token, role, tenant_id)
|
||||
- **useDashboard**: Dashboard data fetching and caching
|
||||
- **useFindings**: Findings list and detail fetching
|
||||
|
||||
### Backend Components
|
||||
|
||||
#### API Routes
|
||||
- **auth.py**: Login, token refresh, user info
|
||||
- **dashboard.py**: Dashboard data aggregation
|
||||
- **findings.py**: Findings CRUD operations
|
||||
- **reports.py**: Audit report generation
|
||||
- **attack_paths.py**: Attack path visualization
|
||||
- **footprint.py**: Digital footprint data
|
||||
- **ai.py**: AI translation and coaching
|
||||
|
||||
#### Services
|
||||
- **ai_translator.py**: OpenAI/Anthropic integration for risk translation
|
||||
- **risk_calculator.py**: Risk score calculation algorithm
|
||||
- **report_generator.py**: PDF report generation with Jinja2/WeasyPrint
|
||||
|
||||
#### Core
|
||||
- **config.py**: Configuration management with Pydantic Settings
|
||||
- **security.py**: JWT token management, password hashing, RBAC decorators
|
||||
|
||||
---
|
||||
|
||||
## Data Model
|
||||
|
||||
### Entity Relationship Diagram
|
||||
|
||||
```
|
||||
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
|
||||
│ tenants │───────│ users │───────│ findings │
|
||||
│─────────────│ 1:N │─────────────│ 1:N │─────────────│
|
||||
│ id (PK) │ │ id (PK) │ │ id (PK) │
|
||||
│ name │ │ tenant_id │ │ tenant_id │
|
||||
│ slug │ │ email │ │ asset_id │
|
||||
│ industry │ │ role │ │ executive_id│
|
||||
│ size_range │ │ ... │ │ severity │
|
||||
│ ... │ └─────────────┘ │ status │
|
||||
└─────────────┘ │ category │
|
||||
│ │ ai_summary │
|
||||
│ │ ... │
|
||||
│ └─────────────┘
|
||||
│ │
|
||||
│ │
|
||||
┌─────────────┐ ┌───────────▼──────────┐
|
||||
│ assets │ │ risk_scores │
|
||||
│─────────────│ │──────────────────────│
|
||||
│ id (PK) │ │ id (PK) │
|
||||
│ tenant_id │ │ tenant_id │
|
||||
│ name │ │ score_date │
|
||||
│ asset_type │ │ overall_score │
|
||||
│ value │ │ score_identity │
|
||||
│ ... │ │ score_cloud │
|
||||
└─────────────┘ │ ... │
|
||||
└──────────────────────┘
|
||||
|
||||
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
|
||||
│ executives │ │authorized │ │attack_paths │
|
||||
│─────────────│ │ assets │ │─────────────│
|
||||
│ id (PK) │ │─────────────│ │ id (PK) │
|
||||
│ tenant_id │ │ id (PK) │ │ finding_id │
|
||||
│ full_name │ │ tenant_id │ │ title │
|
||||
│ title │ │ value │ │ ai_narrative│
|
||||
│ email │ │ asset_type │ │ nodes_json │
|
||||
│ ... │ │ ... │ │ edges_json │
|
||||
└─────────────┘ └─────────────┘ └─────────────┘
|
||||
|
||||
┌─────────────┐
|
||||
│audit_reports│
|
||||
│─────────────│
|
||||
│ id (PK) │
|
||||
│ tenant_id │
|
||||
│ title │
|
||||
│ report_date │
|
||||
│ baseline_ │
|
||||
│ score │
|
||||
│ pdf_path │
|
||||
│ ... │
|
||||
└─────────────┘
|
||||
```
|
||||
|
||||
### Key Tables
|
||||
|
||||
#### tenants
|
||||
Organizational units with complete data isolation.
|
||||
|
||||
- `id`: UUID primary key
|
||||
- `name`: Organization display name
|
||||
- `slug`: URL-friendly identifier (unique)
|
||||
- `industry`: Industry classification
|
||||
- `size_range`: Company size (SMB, mid-market, enterprise)
|
||||
- `contact_email`: Primary contact
|
||||
- `is_active`: Active status for soft deletes
|
||||
|
||||
#### users
|
||||
User accounts with role-based access control.
|
||||
|
||||
- `id`: UUID primary key
|
||||
- `tenant_id`: Foreign key to tenants
|
||||
- `email`: Unique email address
|
||||
- `hashed_password`: Bcrypt hash
|
||||
- `role`: enum (executive, it_admin, trustos_admin)
|
||||
- `is_active`: Account status
|
||||
|
||||
#### findings
|
||||
Security vulnerabilities and exposures.
|
||||
|
||||
- `id`: UUID primary key
|
||||
- `tenant_id`: Foreign key to tenants
|
||||
- `asset_id`: Foreign key to assets (nullable)
|
||||
- `executive_id`: Foreign key to executives (nullable)
|
||||
- `title`: Human-readable title
|
||||
- `severity`: enum (critical, high, medium, low, info)
|
||||
- `status`: enum (open, in_progress, resolved, verified)
|
||||
- `category`: enum (external_exposure, cloud_posture, credential_exposure, etc.)
|
||||
- `technical_description`: Raw technical details
|
||||
- `cve_id`: CVE identifier (if applicable)
|
||||
- `cvss_score`: CVSS score (if applicable)
|
||||
- `ai_summary`: AI-generated plain-English summary
|
||||
- `ai_business_impact`: AI-generated business impact
|
||||
- `ai_remediation_steps`: AI-generated fix steps
|
||||
- `assignee_email`: Assigned team member
|
||||
- `due_date`: Remediation deadline
|
||||
- `is_top_risk`: Flag for top 3 risks
|
||||
|
||||
#### risk_scores
|
||||
Daily snapshots of risk metrics.
|
||||
|
||||
- `id`: UUID primary key
|
||||
- `tenant_id`: Foreign key to tenants
|
||||
- `score_date`: Timestamp of snapshot
|
||||
- `overall_score`: 0-100 overall score
|
||||
- `score_identity`: Identity security score
|
||||
- `score_cloud`: Cloud posture score
|
||||
- `score_network`: Network security score
|
||||
- `score_web`: Web application score
|
||||
- `score_credential`: Credential security score
|
||||
- `score_digital_footprint`: Digital footprint score
|
||||
- `score_third_party`: Third-party risk score
|
||||
- `critical_count`: Count of critical findings
|
||||
- `high_count`: Count of high findings
|
||||
- `medium_count`: Count of medium findings
|
||||
- `low_count`: Count of low findings
|
||||
|
||||
---
|
||||
|
||||
## Authentication & Authorization
|
||||
|
||||
### Authentication Flow
|
||||
|
||||
```
|
||||
1. User submits credentials to POST /api/v1/auth/login
|
||||
↓
|
||||
2. Backend validates credentials against database
|
||||
↓
|
||||
3. Backend generates JWT token with:
|
||||
- sub: user_id
|
||||
- role: user_role
|
||||
- tenant_id: tenant_id
|
||||
- exp: expiration timestamp
|
||||
↓
|
||||
4. Frontend stores token in localStorage
|
||||
↓
|
||||
5. Frontend includes token in Authorization header: Bearer <token>
|
||||
↓
|
||||
6. Backend validates token on each protected request
|
||||
↓
|
||||
7. Backend extracts user context from token
|
||||
↓
|
||||
8. Request proceeds with user context
|
||||
```
|
||||
|
||||
### JWT Token Structure
|
||||
|
||||
```json
|
||||
{
|
||||
"sub": "user-uuid",
|
||||
"role": "executive",
|
||||
"tenant_id": "tenant-uuid",
|
||||
"exp": 1234567890,
|
||||
"iat": 1234567890
|
||||
}
|
||||
```
|
||||
|
||||
### Role-Based Access Control (RBAC)
|
||||
|
||||
Three roles with distinct permissions:
|
||||
|
||||
#### Executive
|
||||
- **Can View**: Dashboard, risk scores, AI summaries, trends
|
||||
- **Cannot View**: Raw CVE data, technical details, other tenants
|
||||
- **Can Modify**: None (read-only)
|
||||
|
||||
#### IT Admin
|
||||
- **Can View**: All findings, technical details, CVEs, remediation steps
|
||||
- **Can Modify**: Finding status, assignees, due dates, resolution notes
|
||||
- **Cannot View**: Other tenants' data
|
||||
|
||||
#### TrustOS Admin
|
||||
- **Can View**: All tenants, all data, system metrics
|
||||
- **Can Modify**: Tenant settings, authorized assets, audit reports
|
||||
- **Can Manage**: Users, roles, system configuration
|
||||
|
||||
### Authorization Middleware
|
||||
|
||||
```python
|
||||
# Example from app/core/security.py
|
||||
def require_executive_or_above(payload: dict = Depends(verify_token)):
|
||||
if payload.get("role") not in ["executive", "it_admin", "trustos_admin"]:
|
||||
raise HTTPException(status_code=403, detail="Insufficient permissions")
|
||||
return payload
|
||||
|
||||
def require_it_or_above(payload: dict = Depends(verify_token)):
|
||||
if payload.get("role") not in ["it_admin", "trustos_admin"]:
|
||||
raise HTTPException(status_code=403, detail="Insufficient permissions")
|
||||
return payload
|
||||
|
||||
def require_admin(payload: dict = Depends(verify_token)):
|
||||
if payload.get("role") != "trustos_admin":
|
||||
raise HTTPException(status_code=403, detail="Admin access required")
|
||||
return payload
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## API Design
|
||||
|
||||
### RESTful Conventions
|
||||
|
||||
TrustOS follows RESTful API design principles:
|
||||
|
||||
- **Resource-Based URLs**: `/api/v1/findings`, `/api/v1/tenants`
|
||||
- **HTTP Methods**: GET (read), POST (create), PATCH (update), DELETE (delete)
|
||||
- **Status Codes**: 200 (success), 201 (created), 400 (bad request), 401 (unauthorized), 403 (forbidden), 404 (not found), 500 (server error)
|
||||
- **JSON Request/Response**: All data is JSON-encoded
|
||||
- **Versioning**: `/api/v1/` prefix for future compatibility
|
||||
|
||||
### API Response Format
|
||||
|
||||
#### Success Response
|
||||
```json
|
||||
{
|
||||
"data": { ... },
|
||||
"meta": {
|
||||
"timestamp": "2024-01-01T00:00:00Z",
|
||||
"request_id": "uuid"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Error Response
|
||||
```json
|
||||
{
|
||||
"error": {
|
||||
"code": "VALIDATION_ERROR",
|
||||
"message": "Invalid input data",
|
||||
"details": { ... }
|
||||
},
|
||||
"meta": {
|
||||
"timestamp": "2024-01-01T00:00:00Z",
|
||||
"request_id": "uuid"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Key API Endpoints
|
||||
|
||||
#### Authentication
|
||||
- `POST /api/v1/auth/login` - Authenticate and receive token
|
||||
- `GET /api/v1/auth/me` - Get current user info
|
||||
|
||||
#### Dashboard
|
||||
- `GET /api/v1/dashboard?tenant_id={id}` - Get dashboard data
|
||||
|
||||
#### Findings
|
||||
- `GET /api/v1/findings?tenant_id={id}` - List findings
|
||||
- `GET /api/v1/findings/{id}` - Get finding detail
|
||||
- `POST /api/v1/findings` - Create finding (IT Admin+)
|
||||
- `PATCH /api/v1/findings/{id}/status` - Update finding status (IT Admin+)
|
||||
|
||||
#### Reports
|
||||
- `GET /api/v1/audit-reports?tenant_id={id}` - List audit reports
|
||||
- `POST /api/v1/audit-reports/generate` - Generate audit report (Admin)
|
||||
|
||||
#### AI
|
||||
- `POST /api/v1/ai/translate/{finding_id}` - Trigger AI translation (IT Admin+)
|
||||
- `GET /api/v1/ai/explain/{finding_id}?question=...` - AI Security Coach (IT Admin+)
|
||||
|
||||
---
|
||||
|
||||
## Frontend Architecture
|
||||
|
||||
### Next.js App Router Structure
|
||||
|
||||
```
|
||||
src/app/
|
||||
├── layout.tsx # Root layout with providers
|
||||
├── page.tsx # Root redirect to /login
|
||||
├── login/
|
||||
│ └── page.tsx # Login page
|
||||
├── dashboard/
|
||||
│ └── page.tsx # Executive dashboard
|
||||
├── findings/
|
||||
│ ├── page.tsx # Findings list
|
||||
│ └── [id]/
|
||||
│ └── page.tsx # Finding detail
|
||||
└── footprint/
|
||||
└── page.tsx # Digital footprint center
|
||||
```
|
||||
|
||||
### State Management
|
||||
|
||||
TrustOS uses React Context for global state:
|
||||
|
||||
```typescript
|
||||
// Auth Context
|
||||
interface AuthContext {
|
||||
token: string | null;
|
||||
role: UserRole | null;
|
||||
tenantId: string | null;
|
||||
name: string | null;
|
||||
login: (email: string, password: string) => Promise<void>;
|
||||
logout: () => void;
|
||||
}
|
||||
```
|
||||
|
||||
### API Client
|
||||
|
||||
Custom API client with token management:
|
||||
|
||||
```typescript
|
||||
const api = {
|
||||
login: (email: string, password: string) =>
|
||||
fetch('/api/v1/auth/login', { method: 'POST', body: ... }),
|
||||
|
||||
dashboard: (tenantId: string) =>
|
||||
fetch(`/api/v1/dashboard?tenant_id=${tenantId}`, {
|
||||
headers: { Authorization: `Bearer ${token}` }
|
||||
}),
|
||||
|
||||
// ... other methods
|
||||
};
|
||||
```
|
||||
|
||||
### Component Design Patterns
|
||||
|
||||
#### Presentational Components
|
||||
- Receive data via props
|
||||
- No side effects
|
||||
- Reusable across contexts
|
||||
|
||||
#### Container Components
|
||||
- Fetch data from API
|
||||
- Manage state
|
||||
- Pass data to presentational components
|
||||
|
||||
#### Higher-Order Components
|
||||
- `withAuth`: Wraps components requiring authentication
|
||||
- `withRole`: Wraps components requiring specific roles
|
||||
|
||||
---
|
||||
|
||||
## Backend Architecture
|
||||
|
||||
### FastAPI Application Structure
|
||||
|
||||
```python
|
||||
# app/main.py
|
||||
from fastapi import FastAPI
|
||||
from fastapi.middleware.cors import CORSMiddleware
|
||||
|
||||
app = FastAPI(
|
||||
title="TrustOS API",
|
||||
version="1.0.0",
|
||||
description="AI-powered cyber resilience platform"
|
||||
)
|
||||
|
||||
# CORS middleware
|
||||
app.add_middleware(
|
||||
CORSMiddleware,
|
||||
allow_origins=["http://localhost:3000"],
|
||||
allow_credentials=True,
|
||||
allow_methods=["*"],
|
||||
allow_headers=["*"],
|
||||
)
|
||||
|
||||
# Include routers
|
||||
app.include_router(auth.router, prefix="/api/v1/auth", tags=["auth"])
|
||||
app.include_router(dashboard.router, prefix="/api/v1/dashboard", tags=["dashboard"])
|
||||
# ... other routers
|
||||
|
||||
@app.on_event("startup")
|
||||
async def startup_event():
|
||||
# Create tables in development
|
||||
if settings.DEBUG:
|
||||
async with engine.begin() as conn:
|
||||
await conn.run_sync(Base.metadata.create_all)
|
||||
|
||||
@app.get("/health")
|
||||
async def health_check():
|
||||
return {"status": "healthy"}
|
||||
```
|
||||
|
||||
### Dependency Injection
|
||||
|
||||
FastAPI's dependency system for authentication and database:
|
||||
|
||||
```python
|
||||
# Database session
|
||||
async def get_db() -> AsyncSession:
|
||||
async with AsyncSessionLocal() as session:
|
||||
try:
|
||||
yield session
|
||||
finally:
|
||||
await session.close()
|
||||
|
||||
# Authentication
|
||||
async def verify_token(authorization: str = Header(...)) -> dict:
|
||||
token = authorization.replace("Bearer ", "")
|
||||
payload = decode_jwt(token)
|
||||
return payload
|
||||
|
||||
# Protected route
|
||||
@router.get("/dashboard")
|
||||
async def get_dashboard(
|
||||
tenant_id: str,
|
||||
payload: dict = Depends(verify_token),
|
||||
db: AsyncSession = Depends(get_db)
|
||||
):
|
||||
# Access payload['role'], payload['tenant_id']
|
||||
# Use db for database operations
|
||||
...
|
||||
```
|
||||
|
||||
### Service Layer Pattern
|
||||
|
||||
Business logic separated from API routes:
|
||||
|
||||
```python
|
||||
# app/services/risk_calculator.py
|
||||
async def recalculate_risk_score(tenant_id: str, db: AsyncSession) -> float:
|
||||
# Business logic for risk calculation
|
||||
findings = await get_open_findings(tenant_id, db)
|
||||
score = calculate_score(findings)
|
||||
await save_risk_snapshot(tenant_id, score, db)
|
||||
return score
|
||||
|
||||
# API route uses service
|
||||
@router.patch("/findings/{id}/status")
|
||||
async def update_status(finding_id: str, body: StatusUpdate, db: AsyncSession = Depends(get_db)):
|
||||
finding = await get_finding(finding_id, db)
|
||||
finding.status = body.status
|
||||
await db.commit()
|
||||
|
||||
# Trigger background risk recalculation
|
||||
asyncio.create_task(recalculate_risk_score(finding.tenant_id))
|
||||
|
||||
return finding
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Database Design
|
||||
|
||||
### Schema Design Principles
|
||||
|
||||
1. **Multi-Tenant by Default**: All tables include `tenant_id`
|
||||
2. **UUID Primary Keys**: Distributed-friendly, no sequence contention
|
||||
3. **Soft Deletes**: `is_active` flags instead of hard deletes
|
||||
4. **Audit Fields**: `created_at`, `updated_at` on all tables
|
||||
5. **JSONB for Flexibility**: Store semi-structured data in JSONB columns
|
||||
|
||||
### Indexing Strategy
|
||||
|
||||
```sql
|
||||
-- Tenant isolation (every query)
|
||||
CREATE INDEX idx_findings_tenant_id ON findings(tenant_id);
|
||||
|
||||
-- Common query patterns
|
||||
CREATE INDEX idx_findings_status ON findings(status);
|
||||
CREATE INDEX idx_findings_severity ON findings(severity);
|
||||
CREATE INDEX idx_findings_category ON findings(category);
|
||||
CREATE INDEX idx_findings_tenant_status ON findings(tenant_id, status);
|
||||
|
||||
-- Time-series queries
|
||||
CREATE INDEX idx_risk_scores_tenant_date ON risk_scores(tenant_id, score_date DESC);
|
||||
|
||||
-- Unique constraints
|
||||
CREATE UNIQUE INDEX idx_users_email ON users(email);
|
||||
CREATE UNIQUE INDEX idx_tenants_slug ON tenants(slug);
|
||||
```
|
||||
|
||||
### Connection Pooling
|
||||
|
||||
```python
|
||||
# SQLAlchemy async engine with connection pooling
|
||||
engine = create_async_engine(
|
||||
settings.DATABASE_URL,
|
||||
echo=False,
|
||||
pool_pre_ping=True, # Verify connections before use
|
||||
pool_size=10, # Base pool size
|
||||
max_overflow=20, # Additional connections under load
|
||||
)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## AI Integration
|
||||
|
||||
### AI Service Architecture
|
||||
|
||||
```
|
||||
┌─────────────┐
|
||||
│ API Route │
|
||||
└──────┬──────┘
|
||||
│
|
||||
┌──────▼──────────┐
|
||||
│ AI Translator │
|
||||
│ Service │
|
||||
└──────┬──────────┘
|
||||
│
|
||||
┌──────▼──────────┐
|
||||
│ LLM Provider │
|
||||
│ (OpenAI/Anthropic)│
|
||||
└─────────────────┘
|
||||
```
|
||||
|
||||
### AI Translation Flow
|
||||
|
||||
```python
|
||||
# 1. Finding created or updated
|
||||
finding = Finding(...)
|
||||
|
||||
# 2. Trigger AI translation (background task)
|
||||
asyncio.create_task(translate_finding_async(finding.id))
|
||||
|
||||
# 3. AI service constructs prompt
|
||||
prompt = f"""
|
||||
Translate this cybersecurity finding:
|
||||
Title: {finding.title}
|
||||
Severity: {finding.severity}
|
||||
CVE ID: {finding.cve_id}
|
||||
Technical Description: {finding.technical_description}
|
||||
"""
|
||||
|
||||
# 4. Call LLM with system prompt
|
||||
system_prompt = """
|
||||
You are TrustOS, an AI cyber resilience advisor.
|
||||
Translate technical findings into plain-English business impact.
|
||||
Output JSON with: summary, business_impact, impact_level, remediation_steps.
|
||||
"""
|
||||
|
||||
# 5. Parse and store AI response
|
||||
data = json.loads(llm_response)
|
||||
finding.ai_summary = data["summary"]
|
||||
finding.ai_business_impact = data["business_impact"]
|
||||
finding.ai_remediation_steps = data["remediation_steps"]
|
||||
await db.commit()
|
||||
```
|
||||
|
||||
### AI Provider Selection
|
||||
|
||||
```python
|
||||
# app/core/config.py
|
||||
AI_PROVIDER = os.getenv("AI_PROVIDER", "openai") # or "anthropic"
|
||||
|
||||
# app/services/ai_translator.py
|
||||
if settings.AI_PROVIDER == "openai":
|
||||
client = AsyncOpenAI(api_key=settings.OPENAI_API_KEY)
|
||||
response = await client.chat.completions.create(
|
||||
model="gpt-4o-mini",
|
||||
messages=[...],
|
||||
response_format={"type": "json_object"}
|
||||
)
|
||||
elif settings.AI_PROVIDER == "anthropic":
|
||||
client = AsyncAnthropic(api_key=settings.ANTHROPIC_API_KEY)
|
||||
response = await client.messages.create(
|
||||
model="claude-3-haiku-20240307",
|
||||
messages=[...]
|
||||
)
|
||||
```
|
||||
|
||||
### AI Caching Strategy
|
||||
|
||||
- **Store AI Responses**: AI translations are stored in the database
|
||||
- **Avoid Redundant Calls**: Check if `ai_summary` exists before re-translating
|
||||
- **Background Processing**: AI calls are async and non-blocking
|
||||
- **Graceful Degradation**: If AI fails, use raw technical description
|
||||
|
||||
---
|
||||
|
||||
## Security Architecture
|
||||
|
||||
### Defense in Depth
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────┐
|
||||
│ 1. Network Security │
|
||||
│ - HTTPS/TLS encryption │
|
||||
│ - CORS restrictions │
|
||||
│ - Rate limiting (planned) │
|
||||
└─────────────────────────────────────────────────────────┘
|
||||
┌─────────────────────────────────────────────────────────┐
|
||||
│ 2. Authentication │
|
||||
│ - JWT token validation │
|
||||
│ - Token expiration (8 hours) │
|
||||
│ - Secure password hashing (bcrypt) │
|
||||
└─────────────────────────────────────────────────────────┘
|
||||
┌─────────────────────────────────────────────────────────┐
|
||||
│ 3. Authorization │
|
||||
│ - Role-based access control │
|
||||
│ - Tenant isolation enforcement │
|
||||
│ - Route-level permission checks │
|
||||
└─────────────────────────────────────────────────────────┘
|
||||
┌─────────────────────────────────────────────────────────┐
|
||||
│ 4. Input Validation │
|
||||
│ - Pydantic schema validation │
|
||||
│ - SQL injection prevention (ORM) │
|
||||
│ - XSS prevention (React escaping) │
|
||||
└─────────────────────────────────────────────────────────┘
|
||||
┌─────────────────────────────────────────────────────────┐
|
||||
│ 5. Data Security │
|
||||
│ - Multi-tenant database isolation │
|
||||
│ - Encrypted secrets management │
|
||||
│ - Audit logging │
|
||||
└─────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### Security Headers
|
||||
|
||||
```python
|
||||
# app/main.py
|
||||
from fastapi.middleware.trustedhost import TrustedHostMiddleware
|
||||
|
||||
app.add_middleware(
|
||||
TrustedHostMiddleware,
|
||||
allowed_hosts=["trustos.com", "*.trustos.com"]
|
||||
)
|
||||
|
||||
# Additional headers via middleware
|
||||
@app.middleware("http")
|
||||
async def add_security_headers(request: Request, call_next):
|
||||
response = await call_next(request)
|
||||
response.headers["X-Content-Type-Options"] = "nosniff"
|
||||
response.headers["X-Frame-Options"] = "DENY"
|
||||
response.headers["X-XSS-Protection"] = "1; mode=block"
|
||||
response.headers["Strict-Transport-Security"] = "max-age=31536000; includeSubDomains"
|
||||
return response
|
||||
```
|
||||
|
||||
### Secrets Management
|
||||
|
||||
- **Environment Variables**: All secrets in `.env` files (git-ignored)
|
||||
- **Production**: Use secret management services (AWS Secrets Manager, HashiCorp Vault)
|
||||
- **Rotation**: Regular API key rotation policy
|
||||
- **Least Privilege**: Database users have minimal required permissions
|
||||
|
||||
---
|
||||
|
||||
## Scalability Considerations
|
||||
|
||||
### Horizontal Scaling
|
||||
|
||||
The architecture supports horizontal scaling:
|
||||
|
||||
- **Stateless API**: FastAPI instances can be scaled horizontally
|
||||
- **Database Connection Pooling**: Efficient connection reuse
|
||||
- **Async I/O**: High concurrency with minimal threads
|
||||
- **Load Balancer**: Nginx or cloud load balancer in front of API
|
||||
|
||||
### Vertical Scaling
|
||||
|
||||
- **Database**: PostgreSQL can scale vertically (more CPU/RAM)
|
||||
- **Caching**: Redis planned for session and query caching
|
||||
- **CDN**: Frontend static assets served via CDN
|
||||
|
||||
### Performance Optimization Strategies
|
||||
|
||||
1. **Database Optimization**
|
||||
- Proper indexing on frequently queried columns
|
||||
- Query optimization with EXPLAIN ANALYZE
|
||||
- Connection pooling to reduce overhead
|
||||
- Read replicas for read-heavy workloads (future)
|
||||
|
||||
2. **API Optimization**
|
||||
- Response compression (gzip)
|
||||
- Pagination for large result sets
|
||||
- Selective field loading (avoid SELECT *)
|
||||
- Async operations throughout
|
||||
|
||||
3. **Frontend Optimization**
|
||||
- Code splitting with Next.js
|
||||
- Image optimization
|
||||
- Static generation where possible
|
||||
- Client-side caching
|
||||
|
||||
4. **Caching Strategy**
|
||||
- API response caching (Redis)
|
||||
- Static asset caching (CDN)
|
||||
- AI response caching (database)
|
||||
- Browser caching headers
|
||||
|
||||
---
|
||||
|
||||
## Performance Optimization
|
||||
|
||||
### Database Query Optimization
|
||||
|
||||
```python
|
||||
# Bad: N+1 query problem
|
||||
for finding in findings:
|
||||
asset = await get_asset(finding.asset_id) # N queries
|
||||
|
||||
# Good: Eager loading
|
||||
findings = await db.execute(
|
||||
select(Finding).options(selectinload(Finding.asset))
|
||||
)
|
||||
```
|
||||
|
||||
### API Response Optimization
|
||||
|
||||
```python
|
||||
# Bad: Return all fields
|
||||
@router.get("/findings")
|
||||
async def list_findings():
|
||||
return await db.execute(select(Finding))
|
||||
|
||||
# Good: Select only needed fields
|
||||
@router.get("/findings")
|
||||
async def list_findings():
|
||||
return await db.execute(
|
||||
select(Finding.id, Finding.title, Finding.severity)
|
||||
)
|
||||
```
|
||||
|
||||
### Frontend Performance
|
||||
|
||||
```typescript
|
||||
// Bad: Re-render on every state change
|
||||
useEffect(() => {
|
||||
fetchData();
|
||||
}, [state]); // Runs on every state change
|
||||
|
||||
// Good: Only re-fetch when dependencies change
|
||||
useEffect(() => {
|
||||
fetchData();
|
||||
}, [tenantId, filter]); // Only runs when tenant or filter changes
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Monitoring & Observability
|
||||
|
||||
### Logging Strategy
|
||||
|
||||
```python
|
||||
# Structured logging
|
||||
import logging
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
logger.info(
|
||||
"Finding status updated",
|
||||
extra={
|
||||
"finding_id": finding.id,
|
||||
"old_status": old_status,
|
||||
"new_status": new_status,
|
||||
"user_id": user_id,
|
||||
"tenant_id": tenant_id
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
### Metrics to Track
|
||||
|
||||
- **API Metrics**: Request rate, error rate, response time
|
||||
- **Database Metrics**: Query time, connection pool usage
|
||||
- **Business Metrics**: Active tenants, findings created, risk score trends
|
||||
- **AI Metrics**: API calls, token usage, cost tracking
|
||||
|
||||
### Health Checks
|
||||
|
||||
```python
|
||||
@app.get("/health")
|
||||
async def health_check():
|
||||
checks = {
|
||||
"database": await check_database(),
|
||||
"ai_service": await check_ai_service(),
|
||||
"storage": await check_storage()
|
||||
}
|
||||
status = "healthy" if all(checks.values()) else "degraded"
|
||||
return {"status": status, "checks": checks}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Future Architecture Enhancements
|
||||
|
||||
### Planned Improvements
|
||||
|
||||
1. **Message Queue**: Celery + Redis for background job processing
|
||||
2. **Caching Layer**: Redis for session and query caching
|
||||
3. **Read Replicas**: PostgreSQL read replicas for scaling reads
|
||||
4. **Microservices**: Split into separate services (auth, findings, monitoring)
|
||||
5. **Event Sourcing**: Event-driven architecture for audit trail
|
||||
6. **GraphQL**: Alternative to REST for complex queries
|
||||
7. **Real-time Updates**: WebSocket for live dashboard updates
|
||||
8. **Edge Computing**: Cloudflare Workers for global distribution
|
||||
|
||||
---
|
||||
|
||||
## Conclusion
|
||||
|
||||
The TrustOS architecture is designed to be secure, scalable, and maintainable while following modern best practices. The multi-tenant isolation, AI augmentation, and authorization-first principles ensure the platform can grow with customer needs while maintaining security and privacy.
|
||||
|
||||
For questions or contributions to the architecture, please refer to the [CONTRIBUTING.md](CONTRIBUTING.md) guide.
|
||||
Reference in New Issue
Block a user