- Enhance README.md with full installation instructions, architecture overview, features, configuration, testing, deployment, security, and troubleshooting sections - Add ARCHITECTURE.md with detailed system architecture, data model, authentication, API design, frontend/backend architecture, database design, AI integration, security, and scalability considerations -Add API.md with complete API reference including authentication, all endpoints, data models, examples, and interactive documentation links - Add DEPLOYMENT.md with deployment guides for Railway, Render, VPS, and Kubernetes, including pre-deployment checklist, monitoring, backup, and troubleshooting - Add CONTRIBUTING.md with development workflow, coding standards, testing guidelines, documentation standards, PR process, and community guidelines
981 lines
35 KiB
Markdown
981 lines
35 KiB
Markdown
# TrustOS Architecture Documentation
|
|
|
|
This document provides a detailed overview of the TrustOS system architecture, design decisions, data flow, and technical implementation details.
|
|
|
|
---
|
|
|
|
## Table of Contents
|
|
|
|
- [System Overview](#system-overview)
|
|
- [Architecture Principles](#architecture-principles)
|
|
- [Component Architecture](#component-architecture)
|
|
- [Data Model](#data-model)
|
|
- [Authentication & Authorization](#authentication--authorization)
|
|
- [API Design](#api-design)
|
|
- [Frontend Architecture](#frontend-architecture)
|
|
- [Backend Architecture](#backend-architecture)
|
|
- [Database Design](#database-design)
|
|
- [AI Integration](#ai-integration)
|
|
- [Security Architecture](#security-architecture)
|
|
- [Scalability Considerations](#scalability-considerations)
|
|
- [Performance Optimization](#performance-optimization)
|
|
|
|
---
|
|
|
|
## System Overview
|
|
|
|
TrustOS is a multi-tenant, AI-powered cyber resilience platform built on a modern microservices-inspired architecture. The system is designed to be:
|
|
|
|
- **Secure**: Multi-tenant isolation with defense-in-depth security
|
|
- **Scalable**: Async I/O throughout for high concurrency
|
|
- **Maintainable**: Clean separation of concerns and modular design
|
|
- **Resilient**: Graceful degradation when external services are unavailable
|
|
|
|
### High-Level Architecture
|
|
|
|
```
|
|
┌─────────────────────────────────────────────────────────────────┐
|
|
│ Client Layer │
|
|
│ Web Browser (Executive, IT Admin, TrustOS Admin) │
|
|
└────────────────────┬────────────────────────────────────────────┘
|
|
│ HTTPS
|
|
┌────────────────────▼────────────────────────────────────────────┐
|
|
│ Frontend Layer │
|
|
│ Next.js 16 + TypeScript + Tailwind CSS + shadcn/ui │
|
|
│ - Server-Side Rendering (SSR) │
|
|
│ - Client-Side Hydration │
|
|
│ - Static Site Generation (SSG) where applicable │
|
|
└────────────────────┬────────────────────────────────────────────┘
|
|
│ REST API (JSON)
|
|
┌────────────────────▼────────────────────────────────────────────┐
|
|
│ API Gateway │
|
|
│ FastAPI Application │
|
|
│ - Request Validation (Pydantic) │
|
|
│ - Authentication (JWT) │
|
|
│ - Authorization (RBAC) │
|
|
│ - Rate Limiting (future) │
|
|
│ - Request Logging │
|
|
└────────────────────┬────────────────────────────────────────────┘
|
|
│
|
|
┌────────────┴────────────┐
|
|
│ │
|
|
┌───────▼────────┐ ┌────────▼─────────┐
|
|
│ Service Layer │ │ Background │
|
|
│ │ │ Workers │
|
|
│ - Dashboard │ │ - AI Translation│
|
|
│ - Findings │ │ - Risk Calc │
|
|
│ - Reports │ │ - PDF Gen │
|
|
│ - Footprint │ │ - Monitoring │
|
|
└───────┬────────┘ └────────┬─────────┘
|
|
│ │
|
|
┌───────▼─────────────────────────▼──────────┐
|
|
│ Data Access Layer │
|
|
│ SQLAlchemy 2.0 (Async ORM) │
|
|
│ - Query Building │
|
|
│ - Connection Pooling │
|
|
│ - Transaction Management │
|
|
└────────────────────┬─────────────────────────┘
|
|
│
|
|
┌────────────────────▼─────────────────────────┐
|
|
│ Database Layer │
|
|
│ PostgreSQL 16 │
|
|
│ - Multi-Tenant Data Isolation │
|
|
│ - Indexing Strategy │
|
|
│ - Foreign Key Constraints │
|
|
│ - JSONB for Flexible Data │
|
|
└──────────────────────────────────────────────┘
|
|
|
|
External Services:
|
|
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
|
|
│ OpenAI API │ │ Anthropic API│ │ HIBP API │
|
|
└──────────────┘ └──────────────┘ └──────────────┘
|
|
```
|
|
|
|
---
|
|
|
|
## Architecture Principles
|
|
|
|
### 1. Authorization First
|
|
|
|
TrustOS only scans and monitors assets that have been explicitly authorized by the tenant organization. This is a core security and privacy principle:
|
|
|
|
- **Authorized Assets Table**: All monitored assets must be pre-registered
|
|
- **Scope Enforcement**: All automated checks respect the authorized asset list
|
|
- **Executive Enrollment**: Executive monitoring requires explicit organizational consent
|
|
- **Audit Trail**: All authorization decisions are logged
|
|
|
|
### 2. Multi-Tenant Isolation
|
|
|
|
Every data access path enforces tenant isolation at multiple layers:
|
|
|
|
- **Database Level**: All tables include `tenant_id` with foreign key constraints
|
|
- **ORM Level**: Queries automatically filter by tenant_id
|
|
- **API Level**: Middleware validates tenant access before processing requests
|
|
- **Application Level**: UI components only display tenant-specific data
|
|
|
|
### 3. Privacy by Design
|
|
|
|
Executive and organizational data is handled with privacy as a foundational requirement:
|
|
|
|
- **Minimal Data Collection**: Only collect data necessary for security assessments
|
|
- **Explicit Consent**: Executive enrollment requires organizational authorization
|
|
- **Data Minimization**: Store only what, not who, where possible
|
|
- **Audit Logging**: All data access is logged for accountability
|
|
|
|
### 4. AI-Augmented, Not AI-Dependent
|
|
|
|
AI features enhance the product but are not required for core functionality:
|
|
|
|
- **Graceful Degradation**: Features work with raw technical data if AI is unavailable
|
|
- **Caching**: AI responses are cached to avoid redundant API calls
|
|
- **Fallback Mechanisms**: System continues operating if AI services are down
|
|
- **Cost Control**: Rate limiting and caching to manage AI API costs
|
|
|
|
### 5. Audit Trail
|
|
|
|
All state changes are tracked with full provenance:
|
|
|
|
- **Timestamps**: `created_at` and `updated_at` on all records
|
|
- **User Attribution**: `created_by` and `updated_by` where applicable
|
|
- **Status Changes**: Finding status transitions are logged with notes
|
|
- **Access Logs**: API requests are logged with user and tenant context
|
|
|
|
---
|
|
|
|
## Component Architecture
|
|
|
|
### Frontend Components
|
|
|
|
#### Page Components
|
|
- **Dashboard Page** (`/dashboard`): Executive view with risk score, top risks, trends
|
|
- **Findings Page** (`/findings`): IT admin view with sortable/filterable table
|
|
- **Finding Detail Page** (`/findings/[id]`): Detailed view with technical and business impact
|
|
- **Login Page** (`/login`): Authentication interface
|
|
- **Footprint Page** (`/footprint`): Digital footprint center
|
|
|
|
#### Reusable Components)
|
|
- **RiskDial**: Circular gauge for cyber health score (0-100)
|
|
- **RiskCard**: Card displaying finding with AI summary and impact
|
|
- **TrendChart**: Line chart for 90-day risk history
|
|
- **RemediationBoard**: Kanban board for finding status tracking
|
|
- **Badge**: Severity and status badges with color coding
|
|
|
|
#### Hooks
|
|
- **useAuth**: Authentication state management (token, role, tenant_id)
|
|
- **useDashboard**: Dashboard data fetching and caching
|
|
- **useFindings**: Findings list and detail fetching
|
|
|
|
### Backend Components
|
|
|
|
#### API Routes
|
|
- **auth.py**: Login, token refresh, user info
|
|
- **dashboard.py**: Dashboard data aggregation
|
|
- **findings.py**: Findings CRUD operations
|
|
- **reports.py**: Audit report generation
|
|
- **attack_paths.py**: Attack path visualization
|
|
- **footprint.py**: Digital footprint data
|
|
- **ai.py**: AI translation and coaching
|
|
|
|
#### Services
|
|
- **ai_translator.py**: OpenAI/Anthropic integration for risk translation
|
|
- **risk_calculator.py**: Risk score calculation algorithm
|
|
- **report_generator.py**: PDF report generation with Jinja2/WeasyPrint
|
|
|
|
#### Core
|
|
- **config.py**: Configuration management with Pydantic Settings
|
|
- **security.py**: JWT token management, password hashing, RBAC decorators
|
|
|
|
---
|
|
|
|
## Data Model
|
|
|
|
### Entity Relationship Diagram
|
|
|
|
```
|
|
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
|
|
│ tenants │───────│ users │───────│ findings │
|
|
│─────────────│ 1:N │─────────────│ 1:N │─────────────│
|
|
│ id (PK) │ │ id (PK) │ │ id (PK) │
|
|
│ name │ │ tenant_id │ │ tenant_id │
|
|
│ slug │ │ email │ │ asset_id │
|
|
│ industry │ │ role │ │ executive_id│
|
|
│ size_range │ │ ... │ │ severity │
|
|
│ ... │ └─────────────┘ │ status │
|
|
└─────────────┘ │ category │
|
|
│ │ ai_summary │
|
|
│ │ ... │
|
|
│ └─────────────┘
|
|
│ │
|
|
│ │
|
|
┌─────────────┐ ┌───────────▼──────────┐
|
|
│ assets │ │ risk_scores │
|
|
│─────────────│ │──────────────────────│
|
|
│ id (PK) │ │ id (PK) │
|
|
│ tenant_id │ │ tenant_id │
|
|
│ name │ │ score_date │
|
|
│ asset_type │ │ overall_score │
|
|
│ value │ │ score_identity │
|
|
│ ... │ │ score_cloud │
|
|
└─────────────┘ │ ... │
|
|
└──────────────────────┘
|
|
|
|
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
|
|
│ executives │ │authorized │ │attack_paths │
|
|
│─────────────│ │ assets │ │─────────────│
|
|
│ id (PK) │ │─────────────│ │ id (PK) │
|
|
│ tenant_id │ │ id (PK) │ │ finding_id │
|
|
│ full_name │ │ tenant_id │ │ title │
|
|
│ title │ │ value │ │ ai_narrative│
|
|
│ email │ │ asset_type │ │ nodes_json │
|
|
│ ... │ │ ... │ │ edges_json │
|
|
└─────────────┘ └─────────────┘ └─────────────┘
|
|
|
|
┌─────────────┐
|
|
│audit_reports│
|
|
│─────────────│
|
|
│ id (PK) │
|
|
│ tenant_id │
|
|
│ title │
|
|
│ report_date │
|
|
│ baseline_ │
|
|
│ score │
|
|
│ pdf_path │
|
|
│ ... │
|
|
└─────────────┘
|
|
```
|
|
|
|
### Key Tables
|
|
|
|
#### tenants
|
|
Organizational units with complete data isolation.
|
|
|
|
- `id`: UUID primary key
|
|
- `name`: Organization display name
|
|
- `slug`: URL-friendly identifier (unique)
|
|
- `industry`: Industry classification
|
|
- `size_range`: Company size (SMB, mid-market, enterprise)
|
|
- `contact_email`: Primary contact
|
|
- `is_active`: Active status for soft deletes
|
|
|
|
#### users
|
|
User accounts with role-based access control.
|
|
|
|
- `id`: UUID primary key
|
|
- `tenant_id`: Foreign key to tenants
|
|
- `email`: Unique email address
|
|
- `hashed_password`: Bcrypt hash
|
|
- `role`: enum (executive, it_admin, trustos_admin)
|
|
- `is_active`: Account status
|
|
|
|
#### findings
|
|
Security vulnerabilities and exposures.
|
|
|
|
- `id`: UUID primary key
|
|
- `tenant_id`: Foreign key to tenants
|
|
- `asset_id`: Foreign key to assets (nullable)
|
|
- `executive_id`: Foreign key to executives (nullable)
|
|
- `title`: Human-readable title
|
|
- `severity`: enum (critical, high, medium, low, info)
|
|
- `status`: enum (open, in_progress, resolved, verified)
|
|
- `category`: enum (external_exposure, cloud_posture, credential_exposure, etc.)
|
|
- `technical_description`: Raw technical details
|
|
- `cve_id`: CVE identifier (if applicable)
|
|
- `cvss_score`: CVSS score (if applicable)
|
|
- `ai_summary`: AI-generated plain-English summary
|
|
- `ai_business_impact`: AI-generated business impact
|
|
- `ai_remediation_steps`: AI-generated fix steps
|
|
- `assignee_email`: Assigned team member
|
|
- `due_date`: Remediation deadline
|
|
- `is_top_risk`: Flag for top 3 risks
|
|
|
|
#### risk_scores
|
|
Daily snapshots of risk metrics.
|
|
|
|
- `id`: UUID primary key
|
|
- `tenant_id`: Foreign key to tenants
|
|
- `score_date`: Timestamp of snapshot
|
|
- `overall_score`: 0-100 overall score
|
|
- `score_identity`: Identity security score
|
|
- `score_cloud`: Cloud posture score
|
|
- `score_network`: Network security score
|
|
- `score_web`: Web application score
|
|
- `score_credential`: Credential security score
|
|
- `score_digital_footprint`: Digital footprint score
|
|
- `score_third_party`: Third-party risk score
|
|
- `critical_count`: Count of critical findings
|
|
- `high_count`: Count of high findings
|
|
- `medium_count`: Count of medium findings
|
|
- `low_count`: Count of low findings
|
|
|
|
---
|
|
|
|
## Authentication & Authorization
|
|
|
|
### Authentication Flow
|
|
|
|
```
|
|
1. User submits credentials to POST /api/v1/auth/login
|
|
↓
|
|
2. Backend validates credentials against database
|
|
↓
|
|
3. Backend generates JWT token with:
|
|
- sub: user_id
|
|
- role: user_role
|
|
- tenant_id: tenant_id
|
|
- exp: expiration timestamp
|
|
↓
|
|
4. Frontend stores token in localStorage
|
|
↓
|
|
5. Frontend includes token in Authorization header: Bearer <token>
|
|
↓
|
|
6. Backend validates token on each protected request
|
|
↓
|
|
7. Backend extracts user context from token
|
|
↓
|
|
8. Request proceeds with user context
|
|
```
|
|
|
|
### JWT Token Structure
|
|
|
|
```json
|
|
{
|
|
"sub": "user-uuid",
|
|
"role": "executive",
|
|
"tenant_id": "tenant-uuid",
|
|
"exp": 1234567890,
|
|
"iat": 1234567890
|
|
}
|
|
```
|
|
|
|
### Role-Based Access Control (RBAC)
|
|
|
|
Three roles with distinct permissions:
|
|
|
|
#### Executive
|
|
- **Can View**: Dashboard, risk scores, AI summaries, trends
|
|
- **Cannot View**: Raw CVE data, technical details, other tenants
|
|
- **Can Modify**: None (read-only)
|
|
|
|
#### IT Admin
|
|
- **Can View**: All findings, technical details, CVEs, remediation steps
|
|
- **Can Modify**: Finding status, assignees, due dates, resolution notes
|
|
- **Cannot View**: Other tenants' data
|
|
|
|
#### TrustOS Admin
|
|
- **Can View**: All tenants, all data, system metrics
|
|
- **Can Modify**: Tenant settings, authorized assets, audit reports
|
|
- **Can Manage**: Users, roles, system configuration
|
|
|
|
### Authorization Middleware
|
|
|
|
```python
|
|
# Example from app/core/security.py
|
|
def require_executive_or_above(payload: dict = Depends(verify_token)):
|
|
if payload.get("role") not in ["executive", "it_admin", "trustos_admin"]:
|
|
raise HTTPException(status_code=403, detail="Insufficient permissions")
|
|
return payload
|
|
|
|
def require_it_or_above(payload: dict = Depends(verify_token)):
|
|
if payload.get("role") not in ["it_admin", "trustos_admin"]:
|
|
raise HTTPException(status_code=403, detail="Insufficient permissions")
|
|
return payload
|
|
|
|
def require_admin(payload: dict = Depends(verify_token)):
|
|
if payload.get("role") != "trustos_admin":
|
|
raise HTTPException(status_code=403, detail="Admin access required")
|
|
return payload
|
|
```
|
|
|
|
---
|
|
|
|
## API Design
|
|
|
|
### RESTful Conventions
|
|
|
|
TrustOS follows RESTful API design principles:
|
|
|
|
- **Resource-Based URLs**: `/api/v1/findings`, `/api/v1/tenants`
|
|
- **HTTP Methods**: GET (read), POST (create), PATCH (update), DELETE (delete)
|
|
- **Status Codes**: 200 (success), 201 (created), 400 (bad request), 401 (unauthorized), 403 (forbidden), 404 (not found), 500 (server error)
|
|
- **JSON Request/Response**: All data is JSON-encoded
|
|
- **Versioning**: `/api/v1/` prefix for future compatibility
|
|
|
|
### API Response Format
|
|
|
|
#### Success Response
|
|
```json
|
|
{
|
|
"data": { ... },
|
|
"meta": {
|
|
"timestamp": "2024-01-01T00:00:00Z",
|
|
"request_id": "uuid"
|
|
}
|
|
}
|
|
```
|
|
|
|
#### Error Response
|
|
```json
|
|
{
|
|
"error": {
|
|
"code": "VALIDATION_ERROR",
|
|
"message": "Invalid input data",
|
|
"details": { ... }
|
|
},
|
|
"meta": {
|
|
"timestamp": "2024-01-01T00:00:00Z",
|
|
"request_id": "uuid"
|
|
}
|
|
}
|
|
```
|
|
|
|
### Key API Endpoints
|
|
|
|
#### Authentication
|
|
- `POST /api/v1/auth/login` - Authenticate and receive token
|
|
- `GET /api/v1/auth/me` - Get current user info
|
|
|
|
#### Dashboard
|
|
- `GET /api/v1/dashboard?tenant_id={id}` - Get dashboard data
|
|
|
|
#### Findings
|
|
- `GET /api/v1/findings?tenant_id={id}` - List findings
|
|
- `GET /api/v1/findings/{id}` - Get finding detail
|
|
- `POST /api/v1/findings` - Create finding (IT Admin+)
|
|
- `PATCH /api/v1/findings/{id}/status` - Update finding status (IT Admin+)
|
|
|
|
#### Reports
|
|
- `GET /api/v1/audit-reports?tenant_id={id}` - List audit reports
|
|
- `POST /api/v1/audit-reports/generate` - Generate audit report (Admin)
|
|
|
|
#### AI
|
|
- `POST /api/v1/ai/translate/{finding_id}` - Trigger AI translation (IT Admin+)
|
|
- `GET /api/v1/ai/explain/{finding_id}?question=...` - AI Security Coach (IT Admin+)
|
|
|
|
---
|
|
|
|
## Frontend Architecture
|
|
|
|
### Next.js App Router Structure
|
|
|
|
```
|
|
src/app/
|
|
├── layout.tsx # Root layout with providers
|
|
├── page.tsx # Root redirect to /login
|
|
├── login/
|
|
│ └── page.tsx # Login page
|
|
├── dashboard/
|
|
│ └── page.tsx # Executive dashboard
|
|
├── findings/
|
|
│ ├── page.tsx # Findings list
|
|
│ └── [id]/
|
|
│ └── page.tsx # Finding detail
|
|
└── footprint/
|
|
└── page.tsx # Digital footprint center
|
|
```
|
|
|
|
### State Management
|
|
|
|
TrustOS uses React Context for global state:
|
|
|
|
```typescript
|
|
// Auth Context
|
|
interface AuthContext {
|
|
token: string | null;
|
|
role: UserRole | null;
|
|
tenantId: string | null;
|
|
name: string | null;
|
|
login: (email: string, password: string) => Promise<void>;
|
|
logout: () => void;
|
|
}
|
|
```
|
|
|
|
### API Client
|
|
|
|
Custom API client with token management:
|
|
|
|
```typescript
|
|
const api = {
|
|
login: (email: string, password: string) =>
|
|
fetch('/api/v1/auth/login', { method: 'POST', body: ... }),
|
|
|
|
dashboard: (tenantId: string) =>
|
|
fetch(`/api/v1/dashboard?tenant_id=${tenantId}`, {
|
|
headers: { Authorization: `Bearer ${token}` }
|
|
}),
|
|
|
|
// ... other methods
|
|
};
|
|
```
|
|
|
|
### Component Design Patterns
|
|
|
|
#### Presentational Components
|
|
- Receive data via props
|
|
- No side effects
|
|
- Reusable across contexts
|
|
|
|
#### Container Components
|
|
- Fetch data from API
|
|
- Manage state
|
|
- Pass data to presentational components
|
|
|
|
#### Higher-Order Components
|
|
- `withAuth`: Wraps components requiring authentication
|
|
- `withRole`: Wraps components requiring specific roles
|
|
|
|
---
|
|
|
|
## Backend Architecture
|
|
|
|
### FastAPI Application Structure
|
|
|
|
```python
|
|
# app/main.py
|
|
from fastapi import FastAPI
|
|
from fastapi.middleware.cors import CORSMiddleware
|
|
|
|
app = FastAPI(
|
|
title="TrustOS API",
|
|
version="1.0.0",
|
|
description="AI-powered cyber resilience platform"
|
|
)
|
|
|
|
# CORS middleware
|
|
app.add_middleware(
|
|
CORSMiddleware,
|
|
allow_origins=["http://localhost:3000"],
|
|
allow_credentials=True,
|
|
allow_methods=["*"],
|
|
allow_headers=["*"],
|
|
)
|
|
|
|
# Include routers
|
|
app.include_router(auth.router, prefix="/api/v1/auth", tags=["auth"])
|
|
app.include_router(dashboard.router, prefix="/api/v1/dashboard", tags=["dashboard"])
|
|
# ... other routers
|
|
|
|
@app.on_event("startup")
|
|
async def startup_event():
|
|
# Create tables in development
|
|
if settings.DEBUG:
|
|
async with engine.begin() as conn:
|
|
await conn.run_sync(Base.metadata.create_all)
|
|
|
|
@app.get("/health")
|
|
async def health_check():
|
|
return {"status": "healthy"}
|
|
```
|
|
|
|
### Dependency Injection
|
|
|
|
FastAPI's dependency system for authentication and database:
|
|
|
|
```python
|
|
# Database session
|
|
async def get_db() -> AsyncSession:
|
|
async with AsyncSessionLocal() as session:
|
|
try:
|
|
yield session
|
|
finally:
|
|
await session.close()
|
|
|
|
# Authentication
|
|
async def verify_token(authorization: str = Header(...)) -> dict:
|
|
token = authorization.replace("Bearer ", "")
|
|
payload = decode_jwt(token)
|
|
return payload
|
|
|
|
# Protected route
|
|
@router.get("/dashboard")
|
|
async def get_dashboard(
|
|
tenant_id: str,
|
|
payload: dict = Depends(verify_token),
|
|
db: AsyncSession = Depends(get_db)
|
|
):
|
|
# Access payload['role'], payload['tenant_id']
|
|
# Use db for database operations
|
|
...
|
|
```
|
|
|
|
### Service Layer Pattern
|
|
|
|
Business logic separated from API routes:
|
|
|
|
```python
|
|
# app/services/risk_calculator.py
|
|
async def recalculate_risk_score(tenant_id: str, db: AsyncSession) -> float:
|
|
# Business logic for risk calculation
|
|
findings = await get_open_findings(tenant_id, db)
|
|
score = calculate_score(findings)
|
|
await save_risk_snapshot(tenant_id, score, db)
|
|
return score
|
|
|
|
# API route uses service
|
|
@router.patch("/findings/{id}/status")
|
|
async def update_status(finding_id: str, body: StatusUpdate, db: AsyncSession = Depends(get_db)):
|
|
finding = await get_finding(finding_id, db)
|
|
finding.status = body.status
|
|
await db.commit()
|
|
|
|
# Trigger background risk recalculation
|
|
asyncio.create_task(recalculate_risk_score(finding.tenant_id))
|
|
|
|
return finding
|
|
```
|
|
|
|
---
|
|
|
|
## Database Design
|
|
|
|
### Schema Design Principles
|
|
|
|
1. **Multi-Tenant by Default**: All tables include `tenant_id`
|
|
2. **UUID Primary Keys**: Distributed-friendly, no sequence contention
|
|
3. **Soft Deletes**: `is_active` flags instead of hard deletes
|
|
4. **Audit Fields**: `created_at`, `updated_at` on all tables
|
|
5. **JSONB for Flexibility**: Store semi-structured data in JSONB columns
|
|
|
|
### Indexing Strategy
|
|
|
|
```sql
|
|
-- Tenant isolation (every query)
|
|
CREATE INDEX idx_findings_tenant_id ON findings(tenant_id);
|
|
|
|
-- Common query patterns
|
|
CREATE INDEX idx_findings_status ON findings(status);
|
|
CREATE INDEX idx_findings_severity ON findings(severity);
|
|
CREATE INDEX idx_findings_category ON findings(category);
|
|
CREATE INDEX idx_findings_tenant_status ON findings(tenant_id, status);
|
|
|
|
-- Time-series queries
|
|
CREATE INDEX idx_risk_scores_tenant_date ON risk_scores(tenant_id, score_date DESC);
|
|
|
|
-- Unique constraints
|
|
CREATE UNIQUE INDEX idx_users_email ON users(email);
|
|
CREATE UNIQUE INDEX idx_tenants_slug ON tenants(slug);
|
|
```
|
|
|
|
### Connection Pooling
|
|
|
|
```python
|
|
# SQLAlchemy async engine with connection pooling
|
|
engine = create_async_engine(
|
|
settings.DATABASE_URL,
|
|
echo=False,
|
|
pool_pre_ping=True, # Verify connections before use
|
|
pool_size=10, # Base pool size
|
|
max_overflow=20, # Additional connections under load
|
|
)
|
|
```
|
|
|
|
---
|
|
|
|
## AI Integration
|
|
|
|
### AI Service Architecture
|
|
|
|
```
|
|
┌─────────────┐
|
|
│ API Route │
|
|
└──────┬──────┘
|
|
│
|
|
┌──────▼──────────┐
|
|
│ AI Translator │
|
|
│ Service │
|
|
└──────┬──────────┘
|
|
│
|
|
┌──────▼──────────┐
|
|
│ LLM Provider │
|
|
│ (OpenAI/Anthropic)│
|
|
└─────────────────┘
|
|
```
|
|
|
|
### AI Translation Flow
|
|
|
|
```python
|
|
# 1. Finding created or updated
|
|
finding = Finding(...)
|
|
|
|
# 2. Trigger AI translation (background task)
|
|
asyncio.create_task(translate_finding_async(finding.id))
|
|
|
|
# 3. AI service constructs prompt
|
|
prompt = f"""
|
|
Translate this cybersecurity finding:
|
|
Title: {finding.title}
|
|
Severity: {finding.severity}
|
|
CVE ID: {finding.cve_id}
|
|
Technical Description: {finding.technical_description}
|
|
"""
|
|
|
|
# 4. Call LLM with system prompt
|
|
system_prompt = """
|
|
You are TrustOS, an AI cyber resilience advisor.
|
|
Translate technical findings into plain-English business impact.
|
|
Output JSON with: summary, business_impact, impact_level, remediation_steps.
|
|
"""
|
|
|
|
# 5. Parse and store AI response
|
|
data = json.loads(llm_response)
|
|
finding.ai_summary = data["summary"]
|
|
finding.ai_business_impact = data["business_impact"]
|
|
finding.ai_remediation_steps = data["remediation_steps"]
|
|
await db.commit()
|
|
```
|
|
|
|
### AI Provider Selection
|
|
|
|
```python
|
|
# app/core/config.py
|
|
AI_PROVIDER = os.getenv("AI_PROVIDER", "openai") # or "anthropic"
|
|
|
|
# app/services/ai_translator.py
|
|
if settings.AI_PROVIDER == "openai":
|
|
client = AsyncOpenAI(api_key=settings.OPENAI_API_KEY)
|
|
response = await client.chat.completions.create(
|
|
model="gpt-4o-mini",
|
|
messages=[...],
|
|
response_format={"type": "json_object"}
|
|
)
|
|
elif settings.AI_PROVIDER == "anthropic":
|
|
client = AsyncAnthropic(api_key=settings.ANTHROPIC_API_KEY)
|
|
response = await client.messages.create(
|
|
model="claude-3-haiku-20240307",
|
|
messages=[...]
|
|
)
|
|
```
|
|
|
|
### AI Caching Strategy
|
|
|
|
- **Store AI Responses**: AI translations are stored in the database
|
|
- **Avoid Redundant Calls**: Check if `ai_summary` exists before re-translating
|
|
- **Background Processing**: AI calls are async and non-blocking
|
|
- **Graceful Degradation**: If AI fails, use raw technical description
|
|
|
|
---
|
|
|
|
## Security Architecture
|
|
|
|
### Defense in Depth
|
|
|
|
```
|
|
┌─────────────────────────────────────────────────────────┐
|
|
│ 1. Network Security │
|
|
│ - HTTPS/TLS encryption │
|
|
│ - CORS restrictions │
|
|
│ - Rate limiting (planned) │
|
|
└─────────────────────────────────────────────────────────┘
|
|
┌─────────────────────────────────────────────────────────┐
|
|
│ 2. Authentication │
|
|
│ - JWT token validation │
|
|
│ - Token expiration (8 hours) │
|
|
│ - Secure password hashing (bcrypt) │
|
|
└─────────────────────────────────────────────────────────┘
|
|
┌─────────────────────────────────────────────────────────┐
|
|
│ 3. Authorization │
|
|
│ - Role-based access control │
|
|
│ - Tenant isolation enforcement │
|
|
│ - Route-level permission checks │
|
|
└─────────────────────────────────────────────────────────┘
|
|
┌─────────────────────────────────────────────────────────┐
|
|
│ 4. Input Validation │
|
|
│ - Pydantic schema validation │
|
|
│ - SQL injection prevention (ORM) │
|
|
│ - XSS prevention (React escaping) │
|
|
└─────────────────────────────────────────────────────────┘
|
|
┌─────────────────────────────────────────────────────────┐
|
|
│ 5. Data Security │
|
|
│ - Multi-tenant database isolation │
|
|
│ - Encrypted secrets management │
|
|
│ - Audit logging │
|
|
└─────────────────────────────────────────────────────────┘
|
|
```
|
|
|
|
### Security Headers
|
|
|
|
```python
|
|
# app/main.py
|
|
from fastapi.middleware.trustedhost import TrustedHostMiddleware
|
|
|
|
app.add_middleware(
|
|
TrustedHostMiddleware,
|
|
allowed_hosts=["trustos.com", "*.trustos.com"]
|
|
)
|
|
|
|
# Additional headers via middleware
|
|
@app.middleware("http")
|
|
async def add_security_headers(request: Request, call_next):
|
|
response = await call_next(request)
|
|
response.headers["X-Content-Type-Options"] = "nosniff"
|
|
response.headers["X-Frame-Options"] = "DENY"
|
|
response.headers["X-XSS-Protection"] = "1; mode=block"
|
|
response.headers["Strict-Transport-Security"] = "max-age=31536000; includeSubDomains"
|
|
return response
|
|
```
|
|
|
|
### Secrets Management
|
|
|
|
- **Environment Variables**: All secrets in `.env` files (git-ignored)
|
|
- **Production**: Use secret management services (AWS Secrets Manager, HashiCorp Vault)
|
|
- **Rotation**: Regular API key rotation policy
|
|
- **Least Privilege**: Database users have minimal required permissions
|
|
|
|
---
|
|
|
|
## Scalability Considerations
|
|
|
|
### Horizontal Scaling
|
|
|
|
The architecture supports horizontal scaling:
|
|
|
|
- **Stateless API**: FastAPI instances can be scaled horizontally
|
|
- **Database Connection Pooling**: Efficient connection reuse
|
|
- **Async I/O**: High concurrency with minimal threads
|
|
- **Load Balancer**: Nginx or cloud load balancer in front of API
|
|
|
|
### Vertical Scaling
|
|
|
|
- **Database**: PostgreSQL can scale vertically (more CPU/RAM)
|
|
- **Caching**: Redis planned for session and query caching
|
|
- **CDN**: Frontend static assets served via CDN
|
|
|
|
### Performance Optimization Strategies
|
|
|
|
1. **Database Optimization**
|
|
- Proper indexing on frequently queried columns
|
|
- Query optimization with EXPLAIN ANALYZE
|
|
- Connection pooling to reduce overhead
|
|
- Read replicas for read-heavy workloads (future)
|
|
|
|
2. **API Optimization**
|
|
- Response compression (gzip)
|
|
- Pagination for large result sets
|
|
- Selective field loading (avoid SELECT *)
|
|
- Async operations throughout
|
|
|
|
3. **Frontend Optimization**
|
|
- Code splitting with Next.js
|
|
- Image optimization
|
|
- Static generation where possible
|
|
- Client-side caching
|
|
|
|
4. **Caching Strategy**
|
|
- API response caching (Redis)
|
|
- Static asset caching (CDN)
|
|
- AI response caching (database)
|
|
- Browser caching headers
|
|
|
|
---
|
|
|
|
## Performance Optimization
|
|
|
|
### Database Query Optimization
|
|
|
|
```python
|
|
# Bad: N+1 query problem
|
|
for finding in findings:
|
|
asset = await get_asset(finding.asset_id) # N queries
|
|
|
|
# Good: Eager loading
|
|
findings = await db.execute(
|
|
select(Finding).options(selectinload(Finding.asset))
|
|
)
|
|
```
|
|
|
|
### API Response Optimization
|
|
|
|
```python
|
|
# Bad: Return all fields
|
|
@router.get("/findings")
|
|
async def list_findings():
|
|
return await db.execute(select(Finding))
|
|
|
|
# Good: Select only needed fields
|
|
@router.get("/findings")
|
|
async def list_findings():
|
|
return await db.execute(
|
|
select(Finding.id, Finding.title, Finding.severity)
|
|
)
|
|
```
|
|
|
|
### Frontend Performance
|
|
|
|
```typescript
|
|
// Bad: Re-render on every state change
|
|
useEffect(() => {
|
|
fetchData();
|
|
}, [state]); // Runs on every state change
|
|
|
|
// Good: Only re-fetch when dependencies change
|
|
useEffect(() => {
|
|
fetchData();
|
|
}, [tenantId, filter]); // Only runs when tenant or filter changes
|
|
```
|
|
|
|
---
|
|
|
|
## Monitoring & Observability
|
|
|
|
### Logging Strategy
|
|
|
|
```python
|
|
# Structured logging
|
|
import logging
|
|
logger = logging.getLogger(__name__)
|
|
|
|
logger.info(
|
|
"Finding status updated",
|
|
extra={
|
|
"finding_id": finding.id,
|
|
"old_status": old_status,
|
|
"new_status": new_status,
|
|
"user_id": user_id,
|
|
"tenant_id": tenant_id
|
|
}
|
|
)
|
|
```
|
|
|
|
### Metrics to Track
|
|
|
|
- **API Metrics**: Request rate, error rate, response time
|
|
- **Database Metrics**: Query time, connection pool usage
|
|
- **Business Metrics**: Active tenants, findings created, risk score trends
|
|
- **AI Metrics**: API calls, token usage, cost tracking
|
|
|
|
### Health Checks
|
|
|
|
```python
|
|
@app.get("/health")
|
|
async def health_check():
|
|
checks = {
|
|
"database": await check_database(),
|
|
"ai_service": await check_ai_service(),
|
|
"storage": await check_storage()
|
|
}
|
|
status = "healthy" if all(checks.values()) else "degraded"
|
|
return {"status": status, "checks": checks}
|
|
```
|
|
|
|
---
|
|
|
|
## Future Architecture Enhancements
|
|
|
|
### Planned Improvements
|
|
|
|
1. **Message Queue**: Celery + Redis for background job processing
|
|
2. **Caching Layer**: Redis for session and query caching
|
|
3. **Read Replicas**: PostgreSQL read replicas for scaling reads
|
|
4. **Microservices**: Split into separate services (auth, findings, monitoring)
|
|
5. **Event Sourcing**: Event-driven architecture for audit trail
|
|
6. **GraphQL**: Alternative to REST for complex queries
|
|
7. **Real-time Updates**: WebSocket for live dashboard updates
|
|
8. **Edge Computing**: Cloudflare Workers for global distribution
|
|
|
|
---
|
|
|
|
## Conclusion
|
|
|
|
The TrustOS architecture is designed to be secure, scalable, and maintainable while following modern best practices. The multi-tenant isolation, AI augmentation, and authorization-first principles ensure the platform can grow with customer needs while maintaining security and privacy.
|
|
|
|
For questions or contributions to the architecture, please refer to the [CONTRIBUTING.md](CONTRIBUTING.md) guide.
|