Architecture Reference
Single-entrypoint reverse proxy
Docker Compose exposes only frontend Nginx on host port 3000. All other services stay on the Docker network. Nginx routes by hostname:
Browser
-> app.localhost:3000
-> React SPA
-> /api/* proxy to api:8080
-> s3.localhost:3000
-> MinIO S3 API at minio:9000
-> minio.localhost:3000
-> optional dev console at minio:9001
The Go API, PostgreSQL, Redis, and MinIO are not browser origins in the default Compose deployment. Frontend code calls /api/v1/... on the same origin, avoiding browser CORS for application API traffic. File previews, screenshots, and report downloads use presigned URLs generated with storage.public_base_url, which defaults to http://s3.localhost:3000 for Compose.
Optional Nginx Basic Auth applies to the app.localhost web app shell. It is not applied to /api/ because hxEASM API requests use Authorization: Bearer ...; requiring Basic Auth there would cause the browser Basic header and the application Bearer header to overwrite each other. API authorization remains enforced by hxEASM JWT/API key authentication and backend RBAC. The s3.localhost route is intentionally not Basic Auth protected by default so presigned file URLs can be opened without another browser auth prompt.
Interactive username/password logins can optionally require two-factor authentication. The second-factor stage is inserted between successful password validation and JWT issuance. API key authentication is not interactive and is not challenged with 2FA. See two-factor-authentication.md.
Docker Compose healthchecks validate operational readiness for frontend, API, worker, PostgreSQL, Redis, and MinIO. API readiness checks PostgreSQL, Redis, and file storage; worker health uses a dedicated subcommand and does not dequeue jobs or run scans. See healthchecks.md.
The proposed high-availability and fault-tolerant deployment architecture is documented separately in high-availability.md. It is a design proposal and is not implemented in the default Docker Compose deployment.
User-driven mutations are recorded in a dedicated append-only Audit Log. It is separate from Exposure Changes: Audit Log records user/admin actions and configuration changes, while Exposure Changes records attack-surface/security state changes. See audit-log.md.
Threats can be discovered by plugins or created manually by authorized operators. Vulnerability is one Threat type. Manual Threats use the common Threat table, status workflow, list/detail API, report pipeline, and exposure-change pipeline where applicable. They are marked with source_plugin=manual and do not use fake scan or plugin IDs.
User avatars are personal profile media, not organization artifacts. Avatar bytes are stored in the configured object storage bucket under user-scoped object keys, while the users table stores only private avatar metadata. Avatars are served through authenticated /api/v1/me/avatar requests and are not listed in the organization-scoped Files page.
Persistent Runtime Storage
Runtime state is stored in host bind mounts under ~/.hxeasm/data, not Docker named volumes. PostgreSQL, Redis, MinIO, and worker-side nuclei state are mounted from this directory. Container deletion, image rebuilds, and docker compose down -v do not delete these host files; explicit host file deletion is required to remove runtime data.
Data model
Core entities
Organization
├── has many: Scopes (pending → approved → rejected)
├── has many: Scans
├── has many: Assets
│ └── has many: AssetEdges (directed graph)
├── has many: Threats (optionally attached to Assets)
├── has many: ExposureChanges (timeline of attack-surface changes)
├── has many: Files (metadata for S3/MinIO objects)
└── has many: Members (User × role)
Asset graph
Assets are stored in assets table, relationships in asset_edges:
assets (id, org_id, type, value, normalized_value, criticality, …)
asset_edges (id, org_id, from_asset_id, to_asset_id, relation_type, …)
Supported current relation types:
owns, contains, resolves_to, points_to, exposes, serves,
has_certificate, has_vulnerability, discovered_by, related_to, generated_candidate
Legacy/historical relation types such as has_path, has_parameter, uses_technology, announces, and discovered_url may exist in older data. New scans do not create those relations as canonical Asset Graph edges.
Canonical EASM graph chain
The backend relation builder keeps attack-surface topology in this order:
organization -> domain -> subdomain -> ip -> service -> webapp -> vulnerability
Allowed branches include:
service/webapp/subdomain/domain -> certificate, preferring service when host/IP/port metadata is available- organization root
ownsedges to root domain/IP Assets that have no stronger discovered parent
HTTP probing is service-first: when a WebApp has IP and port metadata, the worker finds or creates the canonical service asset ip:port/tcp, creates ip -[exposes]-> service, then creates service -[serves]-> webapp. It does not create direct ip -> webapp or subdomain -> webapp edges when a service can be identified. If no IP/service is known, the worker may create a degraded host serves WebApp edge until later DNS/port data enriches the graph.
| Plugin | Output | Required metadata | Created relation |
|---|---|---|---|
subfinder, amass |
subdomain | parent_domain |
domain contains subdomain |
alterx |
candidate subdomain | parent_domain, candidate=true |
parent generated_candidate candidate |
dnsx, resolver, shuffledns |
ip | domain or host, record_type when available |
host resolves_to ip |
naabu, nmap |
service | ip, port, protocol |
ip exposes service |
httpx |
webapp | host, ip, port, scheme when available |
service serves webapp; degraded host serves webapp only without service data |
httpx_screenshot |
file artifact | linked WebApp asset value | no topology edge |
tlsx |
certificate | host, ip, port when available |
preferred source has_certificate certificate |
katana |
webapp metadata | parent_url, WebApp host/port metadata, paths, parameters |
related WebApp edge only when a distinct WebApp is discovered |
nuclei |
vulnerability record | matched_url or target asset metadata |
vulnerability stored against WebApp/service asset |
asnmap |
scan output only | ASN/org scope input | does not create ASN/CIDR Assets |
Technology metadata
Technologies are descriptive metadata, not first-class graph assets. WebApp and service assets can carry metadata.technologies, for example:
{
"technologies": ["nginx", "React", "Cloudflare"],
"title": "Example",
"status_code": 200
}
Paths and parameters are also WebApp metadata, typically stored as metadata.paths and metadata.parameters. This keeps the graph focused on reachable assets and relations while still showing stack details in asset detail panels. Legacy removed Asset rows may remain archived in old databases, but API list and graph responses filter them out by default.
Deduplication
Assets are keyed on (organization_id, type, normalized_value). Normalisation rules:
| Type | Rule |
|---|---|
| domain / subdomain | lowercase, strip trailing dot |
| ip | canonical IP |
| webapp | lowercase, strip default ports, strip trailing slash |
| service | ip:port/protocol |
Manual asset creation uses the same key and returns a conflict if the normalized asset already exists. Manual assets are inventory records with source_plugin=manual; they do not create scan scopes, approve targets, enqueue Redis jobs, or start plugins.
The global scan setting auto_approve_discovered_assets controls only the initial approved value for newly inserted assets created by scan/plugin discovery. API and worker processes read it from PostgreSQL through the same scan settings repository. Approved Scope seed assets and manually created assets keep their explicit approval semantics, and existing assets are not modified when the setting changes or when they are rediscovered.
Vulnerability lifecycle
new → active
new → closed(fp)
active → in_progress
active → closed(skipped)
in_progress → review
review → closed(fixed)
review → active
The canonical workflow stores terminal handling as status=closed plus close_reason=fp|fixed|skipped. The review state is the BPMN "remediation complete, waiting for security-team retest" state.
Asset lifecycle state is stored on assets as approved, disabled, and stale. Existing assets are approved during migration. Approved scope seed assets are created as approved; automatically discovered assets are created as unapproved unless they already exist. Disabled assets remain stored and visible, but are excluded from asset-derived active execution such as manual plugin runs and saved-asset target reconstruction. Stale is manual-only in this version; automatic stale calculation is not implemented.
Asset criticality is a manual business-importance field on assets, with values unknown, low, medium, high, and critical. It defaults to unknown, is updated only by users with admin or hacker role, and is not inferred from scanner output or AI agents in the MVP. It is intended as a foundation for future risk scoring, but no risk-score engine or automatic criticality assignment is implemented here.
Scanner-owned observed technical state is stored under assets.metadata.asset_info. This namespace is separate from assets.managed_context, which is operator-managed. Plugins can emit partial observed-state patches; omitted fields are not removals, and unrelated fields from other plugins are preserved by a dedicated deep merge. Canonical observed-state patches now create full asset_scan_snapshots and patch-aware field-level asset_scan_changes. See Asset Observed State.
Asset History is stored separately in asset_history. It is an append-only per-Asset lifecycle timeline for discovery, manual creation, approval, disabled/stale state, and criticality transitions. It is intentionally separate from audit_logs and exposure_changes: Audit Log records who performed user/admin actions, Exposure Changes records global attack-surface/security changes, and Asset History records what happened to one specific Asset over time.
The Asset History MVP does not backfill existing assets and does not record noisy rediscovery refresh fields such as last_seen_at, updated_at, or full metadata churn.
Files and object storage
Scan-generated files are stored in S3-compatible object storage. Local Docker Compose uses MinIO. PostgreSQL stores metadata in the files table; binary content stays in object storage. API download endpoints return short-lived presigned URLs after organization-scoped RBAC checks.
Plugins return PluginResult.Artifacts with local temporary paths. The worker uploads those artifacts through the Files service and associates them with organization, scan, scan job, source plugin, and asset where possible.
Object key format:
organizations/{org_id}/scans/{scan_id}/{file_type}/{uuid}-{safe_filename}
organizations/{org_id}/files/{file_type}/{uuid}-{safe_filename}
See files.md for storage configuration and API details.
Research Layer
The platform architecture includes a research knowledge layer:
platform code
plugins
configuration
hxresearch knowledge repository
hxresearch/ is intended to become the long-term repository of internal security expertise for the project. It is separate from application code and can contain custom detection templates, proprietary research, advisories, writeups, proof-of-concepts, datasets, and future expert knowledge.
Planned structure:
hxresearch/
├── README.md
├── nuclei/
├── advisories/
├── writeups/
├── poc/
└── datasets/
Worker containers mount hxresearch/ read-only at /opt/hxeasm/hxresearch. The nuclei_custom_templates plugin can run templates from /opt/hxeasm/hxresearch/nuclei when an operator adds that plugin to a custom scan profile. It is not included in built-in profiles.
Future detection capabilities may add a hybrid Nuclei mode, advisory engines, or agent context readers that consume other hxresearch directories.
See hxresearch.md for the canonical documentation.
Scan pipeline
1. User creates scan (POST /api/v1/organizations/{id}/scans)
2. API creates Scan record (status=pending)
3. API pushes QueueMessage to Redis list easm:scan:queue
4. Worker pops message, updates Scan to running
5. For each plugin in profile:
a. Create ScanJob record
b. Execute plugin.Run() (CLI subprocess or mock)
c. Parse NormalizedEntities from result
d. Upload PluginResult.Artifacts to S3/MinIO through Files service
e. Upsert assets to DB (deduplication via ON CONFLICT)
f. Vulnerability entities → Threats table with threat_type=vulnerability
g. Update ScanJob status
6. Update Scan to success / failed
Manual Plugin Execution
Users with admin or hacker role can launch one compatible plugin against one selected asset from the Assets page or Asset Graph. The API endpoint is:
POST /api/v1/assets/{asset_id}/run-plugin
The backend resolves the asset organization, validates RBAC, validates that the plugin is enabled, validates manual_scan support from registry metadata, validates supported asset types, creates a normal scan with internal profile manual_scan, creates exactly one queued scan job, and pushes the existing Redis scan queue. The worker processes it with the normal plugin execution, result persistence, file artifact, graph, vulnerability, and exposure-change paths.
The frontend discovers manual plugins through GET /api/v1/plugins/manual-capabilities; the UI does not keep a separate plugin allowlist. Adding a registered plugin with SupportedExecutionModes: [manual_scan] and SupportedAssetTypes makes it available automatically.
Manual scans differ from normal profile scans only in target selection: the plugin input contains the selected asset only. Before queue execution, manual targets are normalized. Services are converted from stored graph form such as 1.2.3.4:443/tcp to execution form 1.2.3.4:443.
Disabled assets cannot be used for manual plugin execution. Saved disabled assets are also skipped when vuln scans and retries rebuild target sets from the asset inventory. This does not change approved scope behavior: a disabled asset can still be rediscovered if an independent approved scope includes the same infrastructure.
Plugin system
Every scanner implements the Plugin interface:
type Plugin interface {
Name() string
Type() PluginType
Version() string
Run(ctx context.Context, input PluginInput, config PluginConfig) (*PluginResult, error)
}
PluginResult contains []NormalizedEntity — the common schema for all outputs:
{
"entity_type": "asset",
"asset_type": "subdomain",
"value": "api.example.com",
"source_plugin": "subfinder",
"confidence": 0.95,
"metadata": { ... }
}
Vulnerability entities use entity_type: "vulnerability" and carry title, severity, template_id, matched target, and service context in metadata. The vulnerability read API joins the affected asset at request time and exposes asset context fields such as asset_value, asset_type, asset_criticality, affected_host, affected_port, affected_protocol, and affected_service; this avoids duplicating asset data in the vulnerability table.
Core plugin contracts live in backend/internal/plugins: model types, registry, shared command helpers, and result/artifact contracts. Concrete tool integrations live in backend/internal/plugins/wrappers, where contributors add or update wrappers such as httpx, nmap, dnsx, and katana. Default wrapper registration is centralized in backend/internal/plugins/wrappers/defaults.go and used by both API and worker startup.
Technology is metadata, not a graph asset. Plugins should write metadata.technologies on WebApp/service assets; asset persistence normalizes legacy tech and technology keys into that canonical array.
Adding a new plugin
- Create
backend/internal/plugins/wrappers/myplugin.go - Implement
plugins.Plugin - Register in
backend/internal/plugins/wrappers/defaults.go - Set
SupportedAssetTypesandSupportedExecutionModes - Add toolinstaller config if the wrapper calls an external CLI
- Add to relevant
scan_profilesentries inbackend/configs/config.yamlfor profile scans
If SupportedExecutionModes includes manual_scan and the plugin is enabled, it appears in GET /api/v1/plugins/manual-capabilities and in the frontend Run Plugin menu without frontend code changes.
Test mode
Set EASM_TEST_MODE=true. Each plugin checks config.Options["test_mode"] and returns hardcoded mock assets/vulns instead of launching real binaries. Useful for UI development and integration tests.
Queue & retry model
Redis list: easm:scan:queue (LPUSH producer, BRPOP consumer)
Worker uses BRPop with 5 s timeout, runs each scan in a goroutine.
Job statuses: queued → running → success / failed / timeout / cancelled
Retry policy: configured per plugin (PluginConfig.Retry). The worker does not automatically retry on failure. Failed or timed-out scan jobs can be retried manually through POST /api/v1/scan-jobs/{job_id}/retry; the retry creates a new scan job row and runs only the selected plugin using reconstructed organization/scan scope.
API key security
Keys are never stored in plaintext. The generation flow:
- Generate 32 random bytes → hex encode → prefix with
easm_ - SHA-256 hash stored in DB (
key_hash) - First 12 chars stored as
key_prefixfor display - Raw key returned to user once (not stored)
Validation: rehash the provided key, lookup by hash.
RBAC
| Role | Capabilities |
|---|---|
admin |
Full system access, scope approval, user management, all orgs |
hacker |
Assigned orgs only, run scans, triage vulnerabilities |
client |
Read-only: dashboard, approved vulnerabilities, reports |
Role is stored in JWT claims and checked by RequireRole middleware.
Scope seed assets
Scopes remain a separate configuration and scan-input model. Approved organization scopes may seed discovered/managed Assets where that mapping is safe, but the scopes table, status workflow, and scope types are not part of the Asset domain refactor.
When a scope is approved, the backend upserts seed assets with:
{
"source_plugin": "scope",
"confidence": 1.0,
"metadata": {
"source": "scope",
"scope_id": "...",
"seed": true
}
}
Supported seed mappings:
| Scope type | Seed asset behavior |
|---|---|
domain |
Creates a domain asset |
url |
Creates a webapp asset |
ip |
Creates an ip asset |
cidr |
Does not create Assets; scanners may still use the scope input |
ip_range |
Does not create Assets; expanded to individual IP scan targets only at scan execution time |
asn |
Does not create an Asset; scanners may still use the scope input |
org_name |
Does not create an Asset |
CIDR scopes are not expanded by hxEASM in this version. IP range scopes remain one persisted Scope row, but the scan worker expands IPv4 START_IP-END_IP ranges into individual ip scope items before plugin wrappers run. Expansion is capped at 65,536 IPs per range and 65,536 unique expanded range IPs per scan. Range membership alone never creates IP Assets; only actual scanner output can create or update Assets.
Exposure changes
Exposure changes are stored in exposure_changes and provide a read-only timeline of newly observed attack-surface events. The MVP records events when assets, vulnerability Threats, and file artifacts are newly inserted, plus Threat status changes.
The subsystem is best-effort: failures to write exposure changes do not fail scans, Threat updates, or file storage. Duplicate prevention is based on source insert detection and a lightweight service-level similar-event check.
See changes.md for schema, event types, API endpoints, limitations, and roadmap.
Asset History is not derived from Exposure Changes. Some events, such as a new service discovery, may intentionally exist in both systems because one powers the global attack-surface feed and the other powers the per-Asset lifecycle timeline.
Threat Discussions and In-App Notifications
Threat discussions are stored separately from Exposure Changes. Each comment belongs to one Threat and one organization. Backend authorization filters internal comments away from client users before serialization.
Mentions are persisted as resolved user IDs in threat_comment_mentions; raw @username text is not trusted as the notification source of truth. Mention notifications are stored in user_notifications and are scoped to the current user through /api/v1/me/notifications.
The current notification center is REST-based. There are no WebSockets, browser push notifications, or automatic Telegram/email/webhook delivery for vulnerability mentions.
Update Center
The admin Update Center checks the Central Update API and reports whether a newer release is available. Customer hxEASM installations do not call GitHub directly; the private GitHub token is stored only in the Central Update API. It is informational only: the application does not execute update commands, access Docker, or modify local files.
See updates.md for configuration, API endpoints, UI behavior, and manual update commands.
Report Rendering
Report generation uses embedded templates under backend/internal/reports/templates/. HTML reports render the embedded CSS template directly. PDF reports use the same deterministic view model and the existing lightweight native PDF writer for styled sections, tables, findings, and recommendations. JSON and CSV outputs remain available for machine-readable exports.
Asset Criticality
Asset criticality records business importance separately from vulnerability severity. A low-severity issue on a critical VPN may deserve more attention than a medium issue on a test host.
Current behavior is manual-only:
- allowed values:
unknown,low,medium,high,critical - default:
unknown - source: always
manual - update roles:
admin,hacker - read roles: any user with organization access
- audit fields: update time and updating user when available
Future suggested criticality, AI/agent recommendations, and risk-score calculation can build on these fields, but they are not part of the MVP.
Scans Settings and Custom Profiles
Settings -> Scans exposes an admin-only profile and plugin management foundation. Config-backed profiles from backend/configs/config.yaml remain the default source of truth and are read-only in the UI. Admin-created custom profiles are stored in PostgreSQL and merged into the same runtime profile registry as config profiles.
The worker still executes one ordered plugin chain per scan. Before resolving a profile, it refreshes enabled custom profiles from the database, then runs plugins through the existing worker execution path.
Plugin settings are stored separately in plugin_settings and injected into PluginInput.Settings after validation against each plugin's declared option schema. Existing plugins ignore this field unless they explicitly support options.