Skip to content
HolonomiX
Products/HX-SDP
Flagship platform · Private Appliance · Scoped Pilot

HX-SDP

The AI inference data plane as one signed runtime.

HX-SDP combines the functions of twelve services, including vector search, caching, feature storage, event streaming, and gateway operations. It deploys as one signed runtime on customer infrastructure, with GPU-accelerated execution and CPU fallback for scoped evaluation and support validation.

StatusPrivate Appliance
Ship shapesOVA · QCOW2 · OCI · bare-metal
Replaces12 inference-stack services
OperationFull air-gapped operation
LicenseOne runtime · one license
One runtime · One surface · One license · One architecture

Evaluate HX-SDP for your team

Who it serves
AI platform and infrastructure teams responsible for retrieval, serving and data governance.
When to evaluate
Your AI inference stack spans separate storage, retrieval, caching and gateway services, and you need to assess which can share one runtime.
Outcome to establish
A customer-controlled runtime for the agreed data services, with retrieval quality, latency, resource usage and operation receipts evaluated on your workload.
Deployment
Private Appliance · Scoped pilot
Evidence to review
Exact-result and operation receipts, with workload-specific runtime and retrieval results
Commercial starting point
Scope a Private Appliance evaluation or pilot against your existing stack. Agree the corpus, query mix, target quality and latency, infrastructure requirements and baseline operating costs.

The AI inference stack as one platform.

A modern AI inference system is not one thing. It is a stack of services, each from a different vendor, each with its own license, its own operational footprint, its own scaling story, its own integration burden: a vector database, a search index, a feature store, a cache layer, an ETL pipeline, an event stream, an observability stack, a gateway.

Data has structure

One representation throughout inference.

Modern AI infrastructure

Source data

  • Cache
  • Event stream
  • Vector DB
  • Retention
  • Feature store
  • Gateway
  • Search index
  • Observability

Multiple representations of the same signal.

HX-SDP · Structural Data Platform

  • RAG and search
  • Agents
  • Analytics
  • Applications

HX-SDPOne structural representation

  1. 01Classify the structure first
  2. 02Represent it once
  3. 03Preserve that representation throughout inference

A single structural representation throughout inference.

The category is not another vector database.It is a structural data platform for the inference era.

The problem is not any one component. The problem is that nobody designed the system. It accreted. The customer ends up running a federation of vendors and calling it an inference platform.

HX-SDP is the system. One platform replaces the conglomerate. The customer ships one runtime, manages one operational surface, signs one license, scales one set of resources, and debugs one architecture. What used to be a dozen procurement decisions, integration projects, and on-call rotations becomes one.

One platform replaces the conglomerate

The AI inference data plane as one signed runtime.

HX-SDP

GPU-native structural data platform

One runtime · one operational surface · one license

  • Cache
  • Vectors
  • Features
  • Search
  • Retention
  • Gateway
  • Observability

Before serving

Representation policy

Decides the representation before data is served.

Evidence

HX-Provenance

Signs the evidence.

Inference is one system, end to end.

Stack complexity is a choice, not a consequence

The conglomerate exists because the field chose it.

The choice we made

  • CacheLow-latency reads
  • Vector DBSimilarity search
  • Search indexKeyword retrieval
  • Feature storeML features
  • GatewayRouting and governance
  • Event streamReal-time ingestion
  • RetentionPolicies and lifecycle
  • ETLPipelines and transformations

Built independently. Optimized locally. Never revisited the foundational assumption. Complexity compounds when structure is ignored.

The choice we can make

One representation.One runtime.Every access pattern.

Structural awareness, in-place serving.

  • Eliminate hops
  • Reduce cost
  • Increase reliability

Inference can be a system.One representation, held in place, serving every pattern from the same form.

Same signed runtime payload across every shape.

VMware OVA

Enterprise on-prem with vSphere, ESXi, or vCenter-managed clusters with NVIDIA vGPU or PCIe passthrough.

KVM / Proxmox / OpenStack QCOW2

Linux virtualization, private cloud, OpenStack-based federal clouds, Proxmox enterprise environments.

OCI Runtime Kit (Kubernetes)

Containerized enterprise deployments with NVIDIA Container Toolkit, including air-gapped Kubernetes clusters.

Bare-metal Ubuntu installer

Direct hardware install for customers without virtualization, on supported NVIDIA GPU-equipped servers.

Production appliance targets require a supported NVIDIA configuration validated for the supplied release. CPU fallback supports evaluation, first-boot checks and support validation within the agreed package scope; GPU serving performance claims do not apply to CPU-only execution.

Full air-gapped operation supported. License install, validation, and updates work without network access.

Movement is the bill

Data is only expensive when it moves.

Traditional stack

12 service hops. The same data moves again and again.

  1. 01Cache and Vector DB
  2. 02Vector DB and Feature store
  3. 03Cache and ETL
  4. 04Vector DB and ETL
  5. 05Feature store and ETL
  6. 06Cache and Search index
  7. 07Feature store and Gateway
  8. 08ETL and Search index
  9. 09ETL and Gateway
  10. 10ETL to Event stream
  11. 11Retention to ETL
  12. 12Search index and Event stream

Every hop Serialize  Transmit  Stage  Deserialize

In-place serving

No movement across specialized services.

One structural representation

  • Cache
  • Vector DB
  • Feature store
  • ETL
  • Search index
  • Gateway
  • Event stream
  • Retention

The bill collapses with the hops.

Storage is cheap. Movement is the bill.Not faster movement. No movement.

End-to-end HX-SDP latency path

The serving path has ten steps. Steps 03 to 06 run inside HX-SDP.

  1. 01Network, auth and input
  2. 02Query rewrite
  3. HX-SDP serving data plane
    1. 03Structural retrieve
    2. 04Permission filter
    3. 05Rerank and score
    4. 06Top-k assembly
  4. 07Queue
  5. 08Prefill
  6. 09First token
  7. 10Decode stream

HX-SDP advantage

  • Consolidates retrieval stages
  • Eliminates data-plane copies
  • Reduces serving-state overhead
  • Runs as one signed runtime

What changes

  • Retrieval, filtering, reranking and context assembly live inside HX-SDP.
  • Model serving and decoding remain in the inference runtime.
  • One structural representation across the hot path.

Data has structure.HolonomiX structural compute, rendered in one representation.

Performance evidence and operating limits

Match the reported configuration to your corpus, retrieval quality, latency, and infrastructure requirements. Then define the acceptance criteria for your evaluation.

HX-SDP is currently offered through Private Appliance and scoped pilot only. Technical deployment documentation does not establish marketplace availability. Confirm delivery requirements with HolonomiX.

Workload requirements

AI inference workloads that can consolidate structural classification, retrieval, serving, search and governance in one customer-controlled runtime.

Inputs

A corpus and query mix, structural characteristics, target scale, deployment requirements and the data services in scope.

Outputs

Exact-result or operation receipts and release evidence, alongside retrieval quality, latency, resource usage and operating results for the agreed workload.

Limits

Representation and consolidation benefits depend on workload structure and the services in scope.

Availability: Private Appliance · Scoped pilot.