Field note
Split Compute and GPU Runtimes
Keep the Pulsar application and state services together while placing approved inference or creative runtimes on separate GPU hosts.
Evidence status
Implementation path reviewed on July 22, 2026. The review covered external primary-runtime and external ComfyUI placement with the Pulsar application, PostgreSQL, Valkey, and SeaweedFS kept together. It did not benchmark latency, throughput, automatic failover, transport security, or multi-tenant isolation.
Split compute separates GPU execution from the Pulsar application without separating Pulsar’s durable state. The browser remains connected to the application host. Next.js and FastAPI continue to enforce identity, policy, and workflow rules. FastAPI calls an approved external primary runtime for chat, while an ARQ worker can call an approved external ComfyUI service for creative jobs. PostgreSQL, Valkey, and SeaweedFS remain with the Pulsar deployment.
This pattern is useful when a deployment host has enough CPU, memory, and storage for the application but the organization already operates a separate GPU workstation or serving cluster. It can also isolate chat and creative memory budgets. It is not automatically faster, more available, or more secure than a self-contained deployment. Those outcomes depend on the external endpoint, network, identity, capacity, and operating model.
Reviewed topology

Diagram transcript
- 1. The browser connects to customer-managed ingress and the Pulsar Next.js application, never directly to a GPU host.
- 2. Next.js proxies chat to FastAPI, which applies identity, authorization, conversation access, model selection, and policy.
- 3. FastAPI reads durable context from local PostgreSQL, optional managed objects from local SeaweedFS, and applicable admission controls from local Valkey.
- 4. FastAPI sends the synchronous chat request to the approved external primary runtime and streams the response back through Next.js.
- 5. For creative work, FastAPI records a durable job and Valkey with ARQ dispatches it to a local worker.
- 6. The worker calls the approved external ComfyUI endpoint, then stores applicable media in local SeaweedFS and durable status in PostgreSQL.
- 7. Runtime endpoint health, authentication, network protection, model lifecycle, and GPU capacity remain explicit external responsibilities.
Diagram legend
Application and state boundary: Next.js, FastAPI, workers, PostgreSQL, Valkey, and SeaweedFS.
Primary runtime boundary: external Ollama, vLLM, or LM Studio used for synchronous chat.
Creative runtime boundary: optional external ComfyUI used by background workers.
What crosses the boundary
For chat, the application sends the model request required by the selected conversation and tool policy. Depending on the workflow, that request can include prompt text, conversation context, model parameters, and content derived from an approved attachment or tool result. The external runtime returns streamed model output. Pulsar persists the resulting conversation state locally. A runtime endpoint is therefore a data destination that must be included in the deployment’s privacy and egress review.
For creative work, the worker sends the approved generation request and any required reference input to ComfyUI. Generated media returns to the Pulsar workflow and is retained in SeaweedFS when the feature uses managed object storage. Queue and job records remain local. Do not draw a browser-to-ComfyUI path or imply that the external service owns Pulsar users, conversations, approvals, or job history.
| Data or control | Placement in this pattern | Reason |
|---|---|---|
| Users, conversations, jobs, approvals, configuration | Local PostgreSQL | Durable application records remain under the Pulsar deployment owner. |
| Queues, leases, rate limits, transient counters | Local Valkey | Coordination remains close to the application and worker that use it. |
| Managed uploads and generated media | Local SeaweedFS | Object retention and application references remain in the local data boundary. |
| Primary inference execution | Approved external runtime | GPU serving can use separately owned infrastructure. |
| Creative execution | Optional approved external ComfyUI | Image or video GPU work can be separated from chat and application resources. |
Runtime choices and ownership
The external primary endpoint can be an approved Ollama, vLLM, or LM Studio service that presents the compatible API expected by Pulsar. Managed Ollama remains the default for a new self-contained installation, and managed vLLM is the higher-throughput managed option. Choosing an external endpoint changes ownership: Pulsar does not install, upgrade, secure, or guarantee that runtime.
- Name the runtime owner, endpoint class, selected models, authentication method, and maintenance process.
- Verify model discovery, chat completion, streaming, context behavior, and error handling from the Pulsar host.
- Record which data can be sent and whether logs or prompts are retained by the runtime service.
- Separate the optional utility Ollama role from the primary runtime; decide whether it remains local or is omitted.
- Treat external ComfyUI as a distinct creative service with its own models, queue capacity, retention, and change process.
Network and trust requirements
Split compute adds east-west network paths that a self-contained host may not need. Restrict access so only the Pulsar services that require a runtime can reach it. Use the organization’s approved private routing, VPN, proxy, or transport protection. The reviewed application path does not establish built-in mTLS between Pulsar and every external runtime, so transport confidentiality and endpoint authentication must be verified rather than assumed.
Do not expose PostgreSQL, Valkey, or SeaweedFS to the GPU host merely because inference is remote. The runtime should receive a bounded request through its API, not database credentials or storage access. If a creative workflow needs a reference image, pass it through the approved application workflow and document its temporary handling at the external service.
Failure behavior
| Failure | Expected effect | Required response |
|---|---|---|
| Primary runtime unreachable | New chat requests fail or an active stream ends; local durable history remains. | Check routing, endpoint process, model availability, authentication, and capacity. No automatic failover is claimed. |
| ComfyUI unreachable | Creative background jobs fail, wait, or enter workflow-specific recovery while chat can continue. | Inspect durable job state and retry policy before replay; do not clear queues blindly. |
| Network latency or interruption | Streaming may start slowly, stop, or time out; creative transfers may extend job duration. | Measure from the Pulsar host and set an operating threshold based on local evidence. |
| External model changes | Model names, context behavior, output quality, or memory use can change. | Pin and review runtime changes, then rerun representative acceptance checks. |
| Local state service unavailable | Application workflows fail even if GPUs remain healthy. | Recover the owning PostgreSQL, Valkey, or SeaweedFS path locally. |
Verification checklist
- Confirm the browser reaches only the Pulsar ingress and cannot address GPU APIs directly.
- Confirm PostgreSQL, Valkey, and SeaweedFS are private to the application deployment network.
- Test runtime DNS or addressing, transport protection, authentication, model discovery, and streaming from FastAPI.
- Run a representative chat, reload the conversation, and confirm local durable history after a runtime restart.
- Run an approved creative workflow, retrieve the result from Pulsar, and confirm local job and object records.
- Interrupt each external endpoint in a test window and record user-visible errors, timeout behavior, and operator recovery.
- Review external service logs and retention so prompt or media handling matches policy.
- Record capacity and latency as environment-specific observations without turning them into general product claims.
When split compute is not a good fit
Avoid this pattern when the organization cannot protect or operate the network path, cannot identify an owner for the GPU service, or cannot accept request data crossing the application-host boundary. It is also a poor fit when intermittent connectivity would make the primary user workflow unacceptable and no separately designed recovery path exists.
A self-contained deployment may be easier when one host can meet the workload and a compact failure boundary matters more than independent GPU scaling. A managed external platform may be appropriate when the organization already has mature serving controls. Choose from evidence and ownership, not from an assumption that remote GPUs are inherently faster or cheaper.
Common split-compute questions
Does the browser connect to the GPU host?
No. The browser remains connected to Pulsar. FastAPI calls the primary runtime, and the worker calls ComfyUI for queued creative jobs.
Do PostgreSQL and object storage move with the GPU?
No. In the reviewed pattern, PostgreSQL, Valkey, and SeaweedFS remain with the Pulsar application deployment.
Does Pulsar fail over to another runtime automatically?
No automatic runtime failover is established by this field note. A redundant serving design would need separate endpoint, policy, data, and acceptance work.
Apply the pattern deliberately
Use Planning a Pulsar Deployment to compare this topology with a self-contained pattern. Use Operating a Private AI Workspace to add endpoint checks, ownership, incident response, and recovery evidence.
Private deployment consultation
Review Pulsar against your environment.
Bring the infrastructure, security boundaries, model runners, and use cases. The Pulsar team will map the appropriate deployment path.