HX-SDP
The AI inference data plane as one signed runtime.
HX-SDP combines the functions of twelve services, including vector search, caching, feature storage, event streaming, and gateway operations. It deploys as one signed runtime on customer infrastructure, with GPU-accelerated execution and CPU fallback for scoped evaluation and support validation.
Evaluate HX-SDP for your team
- Who it serves
- AI platform and infrastructure teams responsible for retrieval, serving and data governance.
- When to evaluate
- Your AI inference stack spans separate storage, retrieval, caching and gateway services, and you need to assess which can share one runtime.
- Outcome to establish
- A customer-controlled runtime for the agreed data services, with retrieval quality, latency, resource usage and operation receipts evaluated on your workload.
- Deployment
- Private Appliance · Scoped pilot
- Evidence to review
- Exact-result and operation receipts, with workload-specific runtime and retrieval results
- Commercial starting point
- Scope a Private Appliance evaluation or pilot against your existing stack. Agree the corpus, query mix, target quality and latency, infrastructure requirements and baseline operating costs.
The AI inference stack as one platform.
A modern AI inference system is not one thing. It is a stack of services, each from a different vendor, each with its own license, its own operational footprint, its own scaling story, its own integration burden: a vector database, a search index, a feature store, a cache layer, an ETL pipeline, an event stream, an observability stack, a gateway.
Data has structure
One representation throughout inference.
Modern AI infrastructure
Source data
- Cache
- Event stream
- Vector DB
- Retention
- Feature store
- Gateway
- Search index
- Observability
Multiple representations of the same signal.
HX-SDP · Structural Data Platform
- RAG and search
- Agents
- Analytics
- Applications
HX-SDPOne structural representation
- 01Classify the structure first
- 02Represent it once
- 03Preserve that representation throughout inference
A single structural representation throughout inference.
The category is not another vector database.It is a structural data platform for the inference era.
The problem is not any one component. The problem is that nobody designed the system. It accreted. The customer ends up running a federation of vendors and calling it an inference platform.
HX-SDP is the system. One platform replaces the conglomerate. The customer ships one runtime, manages one operational surface, signs one license, scales one set of resources, and debugs one architecture. What used to be a dozen procurement decisions, integration projects, and on-call rotations becomes one.
One platform replaces the conglomerate
The AI inference data plane as one signed runtime.
HX-SDP
GPU-native structural data platform
One runtime · one operational surface · one license
- Cache
- Vectors
- Features
- Search
- Retention
- Gateway
- Observability
Before serving
Representation policy
Decides the representation before data is served.
Evidence
HX-Provenance
Signs the evidence.
Inference is one system, end to end.
Stack complexity is a choice, not a consequence
The conglomerate exists because the field chose it.
The choice we made
- CacheLow-latency reads
- Vector DBSimilarity search
- Search indexKeyword retrieval
- Feature storeML features
- GatewayRouting and governance
- Event streamReal-time ingestion
- RetentionPolicies and lifecycle
- ETLPipelines and transformations
Built independently. Optimized locally. Never revisited the foundational assumption. Complexity compounds when structure is ignored.
The choice we can make
One representation.One runtime.Every access pattern.
Structural awareness, in-place serving.
- Eliminate hops
- Reduce cost
- Increase reliability
Inference can be a system.One representation, held in place, serving every pattern from the same form.
Same signed runtime payload across every shape.
VMware OVA
Enterprise on-prem with vSphere, ESXi, or vCenter-managed clusters with NVIDIA vGPU or PCIe passthrough.
KVM / Proxmox / OpenStack QCOW2
Linux virtualization, private cloud, OpenStack-based federal clouds, Proxmox enterprise environments.
OCI Runtime Kit (Kubernetes)
Containerized enterprise deployments with NVIDIA Container Toolkit, including air-gapped Kubernetes clusters.
Bare-metal Ubuntu installer
Direct hardware install for customers without virtualization, on supported NVIDIA GPU-equipped servers.
Production appliance targets require a supported NVIDIA configuration validated for the supplied release. CPU fallback supports evaluation, first-boot checks and support validation within the agreed package scope; GPU serving performance claims do not apply to CPU-only execution.
Full air-gapped operation supported. License install, validation, and updates work without network access.
Movement is the bill
Data is only expensive when it moves.
Traditional stack
12 service hops. The same data moves again and again.
- 01Cache and Vector DB
- 02Vector DB and Feature store
- 03Cache and ETL
- 04Vector DB and ETL
- 05Feature store and ETL
- 06Cache and Search index
- 07Feature store and Gateway
- 08ETL and Search index
- 09ETL and Gateway
- 10ETL to Event stream
- 11Retention to ETL
- 12Search index and Event stream
Every hop Serialize Transmit Stage Deserialize
In-place serving
No movement across specialized services.
One structural representation
- Cache
- Vector DB
- Feature store
- ETL
- Search index
- Gateway
- Event stream
- Retention
The bill collapses with the hops.
Storage is cheap. Movement is the bill.Not faster movement. No movement.
End-to-end HX-SDP latency path
The serving path has ten steps. Steps 03 to 06 run inside HX-SDP.
- 01Network, auth and input
- 02Query rewrite
- HX-SDP serving data plane
- 03Structural retrieve
- 04Permission filter
- 05Rerank and score
- 06Top-k assembly
- 07Queue
- 08Prefill
- 09First token
- 10Decode stream
HX-SDP advantage
- Consolidates retrieval stages
- Eliminates data-plane copies
- Reduces serving-state overhead
- Runs as one signed runtime
What changes
- Retrieval, filtering, reranking and context assembly live inside HX-SDP.
- Model serving and decoding remain in the inference runtime.
- One structural representation across the hot path.
Data has structure.HolonomiX structural compute, rendered in one representation.
Performance evidence and operating limits
Match the reported configuration to your corpus, retrieval quality, latency, and infrastructure requirements. Then define the acceptance criteria for your evaluation.
HX-SDP is currently offered through Private Appliance and scoped pilot only. Technical deployment documentation does not establish marketplace availability. Confirm delivery requirements with HolonomiX.
Workload requirements
AI inference workloads that can consolidate structural classification, retrieval, serving, search and governance in one customer-controlled runtime.
Inputs
A corpus and query mix, structural characteristics, target scale, deployment requirements and the data services in scope.
Outputs
Exact-result or operation receipts and release evidence, alongside retrieval quality, latency, resource usage and operating results for the agreed workload.
Limits
Representation and consolidation benefits depend on workload structure and the services in scope.
Availability: Private Appliance · Scoped pilot.