Architecture guide
Pulsar Architecture Guide
A practical system model for Pulsar application services, durable state, transient coordination, AI runtimes, tools, and external trust boundaries.
Pulsar is a private AI workspace composed of an application plane, a durable state plane, a transient coordination plane, model runtimes, controlled tools, and optional external services. The design keeps those responsibilities distinct so operators can reason about data ownership, failures, and approvals. A browser does not connect directly to a database or GPU service. Requests enter through the deployment’s ingress, pass through the Pulsar application, and reach only the services required for that workflow.
The most important distinction is between interactive and background work. Chat is synchronous: Browser -> Next.js -> FastAPI -> selected model runtime, with the response streamed back to the browser. Chat never enters ARQ. Media generation, schedules, retention work, recovery work, and enabled Pulsar Code jobs can use durable records plus queued background execution. Conflating those paths produces the wrong capacity plan and the wrong failure assumptions.
System context and trust boundaries
A useful architecture review starts at trust boundaries, not container names. Pulsar expects customer-managed ingress and transport policy in front of the application. The application boundary contains Next.js, FastAPI, and background workers. The private state boundary contains PostgreSQL, Valkey, and SeaweedFS. The runtime boundary contains the selected primary model server and any separate utility or creative runtime. Tool and external-service boundaries contain destinations that may receive explicitly approved requests.

Diagram transcript
- 1. A browser, API client, or operator enters through customer-managed ingress and transport controls.
- 2. Next.js serves the workspace and proxies same-origin API and streaming requests to FastAPI.
- 3. FastAPI applies identity, authorization, policy, and workflow validation before reaching state, runtimes, or tools.
- 4. PostgreSQL, Valkey, and SeaweedFS remain private deployment services with different ownership duties.
- 5. FastAPI calls the selected primary runtime directly for synchronous chat; background workers run queued workflows.
- 6. Optional utility, creative, Pulsar Code, MCP, and external services are reached only by enabled, policy-controlled paths.
Diagram legend
Application plane: Next.js, FastAPI, and ARQ workers.
State plane: PostgreSQL durable records and SeaweedFS managed object data.
Coordination plane: Valkey queues, leases, rate limits, and transient counters.
Runtime plane: managed or external chat, utility, image, and video runtimes.
External boundary: approved services outside the Pulsar deployment.
| Boundary | Primary responsibility | What should not be assumed |
|---|---|---|
| Customer ingress | TLS termination, exposure policy, upstream routing, and any organization-specific edge controls. | Pulsar does not imply a bundled public reverse proxy or a universal network design. |
| Application plane | User experience, API policy, authorization, orchestration, streaming, and background execution. | Application containers are not the durable source of truth. |
| Private state plane | Durable relational records and managed object data. | These services should not be exposed directly to end users or public networks. |
| Coordination plane | Queues, leases, rate limits, wakeups, and short-lived counters. | Transient coordination is not a replacement for durable job or approval records. |
| Runtime plane | Primary inference, optional utility work, and optional creative generation. | A reachable runtime is not automatically trusted, healthy, or sized for the workload. |
| Tools and external services | Explicitly enabled capabilities such as retrieval, repository access, or external model endpoints. | Enabling a tool does not remove the need for destination policy, credentials, and audit review. |
The synchronous chat path
A chat turn stays on the request path because the user is waiting for tokens. The browser sends the turn to a same-origin Next.js route. That proxy preserves the streaming response while FastAPI authenticates the caller, validates the conversation and selected model, obtains the required context, and enforces policy. FastAPI then calls the active OpenAI-compatible model runtime directly and relays server-sent events to the browser.
PostgreSQL provides durable conversation, message, user, model-selection, and policy records. SeaweedFS may provide an attachment or managed object needed by the turn. Valkey can enforce a rate limit, concurrency counter, or stream lease. These supporting reads do not turn chat into a queued job. ARQ is deliberately absent from the chat sequence, so worker backlog does not become the mechanism that schedules interactive tokens.

Diagram transcript
- 1. The browser sends a chat turn to the same-origin Next.js streaming route.
- 2. Next.js proxies the request to FastAPI and keeps the streaming connection open.
- 3. FastAPI authenticates the user, checks conversation access, validates the model, and applies policy.
- 4. FastAPI loads durable context from PostgreSQL and any explicitly referenced object from SeaweedFS.
- 5. FastAPI uses Valkey for applicable admission controls such as rate limits, concurrency, and a stream lease.
- 6. FastAPI calls the selected model runtime directly and relays server-sent events through Next.js to the browser.
- 7. Usable output and final status are persisted, and transient coordination is released or allowed to expire safely.
Diagram legend
Solid numbered steps describe request order; the text order remains authoritative on small screens.
Durable context means PostgreSQL records and only the managed objects referenced by the turn.
Selected runtime can be managed Ollama, managed vLLM, or an approved external compatible endpoint.
Architecture assertion: chat is never an ARQ job
If an architecture diagram, runbook, or capacity model places ARQ between FastAPI and the chat runtime, treat it as obsolete. ARQ workers serve durable background workflows. They do not broker the live chat token stream.
Durable background workflows
Background work begins with a durable application decision. FastAPI validates the request and records a job, schedule, or task in PostgreSQL before dispatch. Valkey and ARQ provide queue delivery, wakeups, leases, or coordination. A worker loads the durable record, executes the permitted work, and writes progress, results, events, and terminal state back to PostgreSQL. When the result is object-backed media, the worker writes the object to SeaweedFS and records its reference.
This pattern supports retries and operator inspection without pretending that a queue item is the business record. Pulsar Code is the clearest example: PostgreSQL remains the authoritative job ledger, while Valkey wakes a worker. Media and scheduler workflows also use background execution, but their exact recovery rules differ. Operators should inspect the workflow-specific status rather than assume every queue has identical replay semantics.

Diagram transcript
- 1. A user or schedule submits a background action to FastAPI.
- 2. FastAPI checks authorization, policy, inputs, and any required approval.
- 3. FastAPI creates the durable job or run record in PostgreSQL.
- 4. Valkey and ARQ dispatch an identifier or wake-up signal to a worker.
- 5. The worker loads authoritative state and calls the allowed scheduler, creative runtime, retention handler, recovery action, or Pulsar Code runner.
- 6. Applicable media objects are stored in SeaweedFS; durable status, events, and references are stored in PostgreSQL.
- 7. Retry, quarantine, failure, or completion follows the workflow-specific policy and remains inspectable by an operator.
Diagram legend
PostgreSQL is the durable ledger before and after execution.
Valkey is transient coordination and queue infrastructure.
SeaweedFS appears only when a workflow creates or consumes managed object data.
Data ownership by service
PostgreSQL is the durable system of record
PostgreSQL owns durable application records: users, roles, conversations, messages, saved configuration, jobs, schedules, approvals, audit-oriented events, and references to managed objects. A restart of an application container or a loss of transient coordination must not redefine those records. Backup and restore planning therefore begins with PostgreSQL and includes consistency with object references.
Valkey coordinates transient work
Valkey is an open-source Redis-compatible data store. Pulsar uses Valkey for queues, rate limits, concurrency controls, leases, wakeups, and ephemeral coordination. Valkey supports the Redis-compatible protocol used by Pulsar application clients. Compatibility identifiers or connection schemes may retain the word redis at an internal protocol boundary; the bundled service remains Valkey.
Treat Valkey data according to the feature that created it. A lost rate-limit counter has a different effect from a delayed background wakeup. Durable status should be reconstructed from PostgreSQL where the workflow provides that recovery behavior. Do not infer that every transient key can be discarded during an incident without impact; use the operating procedure for the affected queue or lease.
SeaweedFS stores managed object data
SeaweedFS provides the bundled S3-compatible object layer for data Pulsar explicitly manages as objects, including applicable uploaded files, generated images, generated video, and other retained media. PostgreSQL stores the application record and object reference. SeaweedFS is not the destination for every export: some generated documents or responses can stream directly to the requester, and some features can use a different storage path by design.
Object backups must be coordinated with the relational records that point to them. Restoring only PostgreSQL can leave missing objects; restoring only object data can leave unreferenced files. A production runbook should state the backup order, retention policy, verification method, and restoration test for both sides without assuming the bundled topology is highly available.
Model runtime roles
Pulsar separates the primary chat runtime from optional utility and creative runtimes. That separation matters because the primary runtime is on the latency-sensitive chat path, while utility tasks can be smaller and creative models have different memory and queue behavior. A single GPU may serve more than one role only when the operator accepts model loading, reduced concurrency, and memory contention.
| Runtime path | Architectural role | Planning note |
|---|---|---|
| Managed Ollama | Default primary runtime for a new managed installation. | A practical starting point for local model serving; installer validation calibrates fit to the host. |
| Managed vLLM | Higher-throughput primary runtime option for suitable GPU deployments. | Choose when supported hardware and workload justify a dedicated serving stack. |
| External Ollama, vLLM, or LM Studio | Primary runtime outside the Pulsar deployment boundary. | The operator owns reachability, endpoint protection, model availability, capacity, and lifecycle. |
| Utility Ollama | Optional separate runtime for bounded utility work such as titles, summaries, or diagnostics. | It is distinct from the primary chat runtime and should be sized and governed separately. |
| ComfyUI | Optional creative runtime for approved image, edit, or video workflows. | Can be local or external; queued jobs and generated object data need explicit capacity and retention plans. |
All primary chat paths use an OpenAI-compatible interface at the application boundary, but operational equivalence should not be assumed. Authentication options, model naming, context behavior, streaming details, and health visibility vary by runtime. Review runtime integration guidance and verify the selected endpoint with the actual model and workload before declaring readiness.
Tools, MCP, and external services
Tools extend a model turn beyond text generation. A retrieval tool may fetch an approved URL. A repository tool may inspect code or prepare a bounded change. An MCP server may expose a domain-specific capability. Pulsar places tool availability and access policy in the application plane so the selected model does not gain ambient network or repository authority merely because it can request a tool.
Every enabled external destination creates a separate trust decision. Operators should record what data can leave the deployment, which identity is used, whether the destination is customer-managed, how credentials rotate, and what logs are retained. A self-hosted application can still send data to an external runtime or tool when configured to do so. Therefore, self-hosted should describe deployment ownership, not be treated as a universal zero-egress claim.
Deployment patterns
The core application and private state services remain deployment-local across the supported patterns. What changes is mainly the placement of AI runtimes and approved tools. A self-contained pattern keeps the primary runtime on the Pulsar host or trusted local infrastructure. An external-runtime pattern sends inference to an approved endpoint. A hybrid pattern keeps primary chat local while sending creative jobs to another GPU host, or uses another deliberate combination.
- Self-contained: simpler ownership and fewer network dependencies, but application services and model workloads compete for host resources unless separated locally.
- External runtime: independent GPU scaling and lifecycle, but additional network, identity, endpoint, data-transfer, and availability responsibilities.
- Hybrid: targeted separation for chat, utility, or creative work, but more failure modes and more configuration to verify.
Use Planning a Pulsar Deployment for the decision process and Split Compute and GPU Runtimes for a bounded topology example. External does not mean that PostgreSQL, Valkey, or SeaweedFS should be moved outside the Pulsar deployment without a separate design review.
Failure boundaries and recovery expectations
| Failure | Likely effect | Operator response |
|---|---|---|
| Primary runtime unavailable | New chat turns fail or stop streaming; durable conversation history remains. | Confirm endpoint health, credentials, model availability, and capacity. Pulsar does not claim automatic runtime failover. |
| ARQ worker unavailable | Queued background work waits while synchronous chat may continue. | Inspect worker health and queue state, then use the approved restart or recovery procedure. |
| Valkey unavailable | Admission controls, leases, and queued dispatch are impaired. | Protect durable state, restore coordination service, and inspect workflow-specific recovery before replay. |
| PostgreSQL unavailable | Durable application operations cannot proceed safely. | Treat as a primary service incident; restore database health before resuming writes. |
| SeaweedFS unavailable | Object-backed uploads or generated media may fail while non-object paths can differ. | Restore object access and verify consistency with PostgreSQL references. |
| External tool unavailable | Only workflows requiring that destination should fail or degrade. | Disable or repair the integration; do not broaden credentials or bypass policy as a recovery shortcut. |
The architecture supports observable, controlled recovery; it does not promise autonomous remediation. The operating surface can analyze redacted evidence and offer a small allowlist of safe actions, but an operator initiates those actions. Restarts, queue clearing, configuration changes, data deletion, and repository publication remain explicit decisions. See Operating a Private AI Workspace for the exact boundary.
Architecture review checklist
- Identify the customer-managed ingress, TLS, identity, and network exposure controls.
- Keep Next.js and FastAPI as the user-facing application boundary; do not expose state services directly.
- Confirm that chat reaches the selected runtime synchronously and does not enter ARQ.
- Map each durable record to PostgreSQL and each managed object to an explicit SeaweedFS-backed feature.
- Map Valkey use to a queue, lease, rate limit, counter, or other transient coordination purpose.
- Name the primary, utility, and creative runtimes separately, including who owns each endpoint.
- List every enabled tool or external service and the data, identity, credentials, and logs it can receive.
- Document backup consistency, restore tests, capacity checks, and approval boundaries without assuming high availability.
Common architecture questions
Does Pulsar send chat through a background queue?
No. Interactive chat follows Browser -> Next.js -> FastAPI -> selected model runtime and streams the result back over the open request. ARQ is reserved for background workflows.
What is the durable source of truth?
PostgreSQL is the durable application system of record. SeaweedFS stores managed object data for features that use the object layer. Valkey provides transient coordination and queue infrastructure rather than replacing durable records.
Can Pulsar use a model runtime on another host?
Yes. Approved external Ollama, vLLM, and LM Studio endpoints are supported runtime paths, and ComfyUI can also be external. The operator owns secure connectivity, endpoint lifecycle, capacity, and data-transfer policy.
Does self-hosted mean no data can leave the deployment?
No. Self-hosted describes ownership of the Pulsar application deployment. Configured external runtimes, tools, or services can receive request data. Egress policy must be evaluated from the enabled configuration.
Is high availability built in?
The reviewed bundled topology does not establish a high-availability guarantee. Organizations that require redundant ingress, replicated state, runtime failover, or tested recovery objectives must design, operate, and verify those controls for their environment.
Next decision
Once the boundaries and request paths are accepted, continue to Planning a Pulsar Deployment. That guide converts this system model into choices about hardware, runtime placement, storage, networking, verification, and ownership.
Private deployment consultation
Review Pulsar against your environment.
Bring the infrastructure, security boundaries, model runners, and use cases. The Pulsar team will map the appropriate deployment path.