Edge & GPU Computing
Edge compute at Dollu PoPs and GPU instances for training and inference - so latency-critical applications and private AI for voice, contact-centre and analytics workloads run on your carrier's network, not across the internet.

40+
PoPs eligible for edge nodes
< 10 ms
Edge round-trip in metro
H100 / L40S
GPU classes available
What is Edge & GPU Computing?
Dollu Edge & GPU Computing places compute where distance matters. Edge nodes at Dollu PoPs and regional data centres run your containers and virtual machines within a few milliseconds of end users, branch sites, private 5G radios and the Dollu voice and messaging core, while GPU instances in our Mumbai, Frankfurt, London, Ashburn and Singapore regions provide the acceleration needed for model training, fine-tuning and high-throughput inference. Both are provisioned and operated with the same portal, API, Terraform provider and 24×7 NOC as the rest of Dollu Cloud.
Edge nodes are built for real-time work: local breakout for private 5G and SD-WAN traffic, media processing close to the caller, IoT ingestion and filtering, video analytics and application caching. Because they sit on the Dollu backbone, they can be stitched into your MPLS or SD-WAN VRF, receive traffic from Dollu DIA or Ethernet services and reach central regions over private paths. Nodes are deployed as managed Kubernetes clusters or VM hosts, from single-server footprints to multi-rack, in Dollu PoPs, colocation sites or on your premises.
GPU instances range from single L40S cards for inference and media workloads to multi-GPU H100 nodes with NVLink and high-bandwidth interconnect for distributed training. Instances come with pre-built images for CUDA, PyTorch, TensorFlow, Triton Inference Server and vLLM, NVMe scratch storage, high-throughput object storage for datasets and checkpoints, and optional managed Kubernetes with GPU scheduling. Capacity is available on demand, reserved or as dedicated bare-metal GPU servers.
Private AI is where the two meet. Speech recognition, speaker analytics, real-time translation, agent assist, call summarisation, fraud detection and conversational IVR all depend on low-latency access to live media and on keeping call content in-country and out of third-party APIs. Dollu runs these models on GPU capacity next to the voice core and at the edge, feeds them from Dollu CCaaS, SIP trunking and CPaaS media streams over private paths, and lets you bring your own models or use curated open-weight models under your control.
Why choose Dollu for edge & gpu.
The advantages of buying from a carrier that owns its network, interconnects and operations - rather than a reseller.
Latency you can design around
Metro round-trips under 10 milliseconds and deterministic private paths make real-time media processing, industrial control loops and interactive applications feasible without central-cloud delays.
AI on your data, in your jurisdiction
Models run on GPUs in the region you choose, fed over private paths, so call recordings, transcripts and customer data never leave the country or pass through a third-party API.
Next to the voice core
Speech and conversational AI workloads receive live media from Dollu SIP, CCaaS and CPaaS services with single-digit-millisecond latency, which makes real-time agent assist and translation practical.
GPU capacity without a capital project
On-demand and reserved GPU instances, or dedicated bare-metal GPU servers, with pre-built software stacks - no data-hall build, power upgrade or hardware procurement cycle.
One network, one operator
Edge nodes, GPUs, private 5G, SD-WAN and voice services from a single provider, monitored by one NOC, so the network path is engineered as part of the application rather than assumed.
Capabilities in detail.
Everything included with Edge & GPU Computing - the platform features, options and controls you get from day one.
- 01
Edge node deployment options
Managed edge clusters at Dollu PoPs, in partner colocation, or on-premises hardware we ship and operate; footprints from a single 1U server to multi-rack; Kubernetes or VM hosts.
- 02
Network integration
Edge nodes stitched into MPLS/SD-WAN VRFs, local breakout for private 5G user-plane traffic, anycast and GSLB steering to the nearest healthy node, private backhaul to Dollu regions.
- 03
GPU instance families
L40S instances for inference, media and graphics; H100 SXM nodes with NVLink and 400G interconnect for training; fractional GPU options for development; dedicated bare-metal GPU servers.
- 04
AI software stack
Images with CUDA, cuDNN, PyTorch, TensorFlow, Triton Inference Server, vLLM and NVIDIA NIM-compatible containers; managed Kubernetes with GPU operator, node pools and autoscaling.
- 05
Data and storage
Local NVMe scratch, high-throughput object storage for datasets and checkpoints, parallel file systems for multi-node training, and private paths from your data sources.
- 06
Private AI for voice and contact centre
Speech-to-text, real-time translation, agent assist, call summarisation, sentiment and compliance monitoring, and conversational IVR, deployed as services fed by Dollu voice media streams.
- 07
Model lifecycle and MLOps
Fine-tuning pipelines, model registry, A/B and canary serving, GPU utilisation monitoring and cost allocation per model or team; bring your own models or use curated open-weight models.
- 08
Security and isolation
Dedicated tenancy options, encrypted storage with customer-managed keys, private endpoints, no telemetry to third parties, and audit logs of every inference and training job.
How it works.
From first conversation to live traffic - a tracked, engineer-led onboarding with a named owner at every step.
- Step 01
Define
We capture latency targets, data locations, throughput and model requirements, and identify which PoPs, regions and GPU classes fit the workload.
- Step 02
Provision
Edge nodes are deployed at the selected PoPs or shipped to your sites; GPU instances or bare-metal servers are provisioned in-region with the software stack you selected.
- Step 03
Connect
Nodes and GPUs are attached to your VRF, private 5G core, voice media streams or data sources over private paths, with steering and failover configured.
- Step 04
Deploy and tune
Your team or ours deploys models and applications, validates latency and throughput against targets, and tunes batch sizes, quantisation and scheduling for cost.
- Step 05
Operate
24×7 NOC monitoring of nodes, GPUs and network paths, utilisation and cost reporting, and capacity planning as usage grows.
Who uses this and why.
Typical deployments across carriers, enterprises, platforms and contact centres.
Contact centres and CCaaS
Real-time transcription, agent assist, translation and post-call summarisation running on GPUs beside the voice core, keeping call content in-country and latency low enough to matter mid-call.
Conversational AI and voicebots
Speech recognition, LLM reasoning and text-to-speech served from GPUs with sub-second turn-taking, connected directly to Dollu SIP trunks and CPaaS voice APIs.
Video analytics and smart sites
Edge nodes at or near campuses, ports and factories process camera streams locally for safety, quality and security, sending only events and metadata upstream.
Manufacturing and private 5G
Local breakout and edge compute for AGVs, machine vision and control systems on private 5G, with deterministic latency and no dependence on internet backhaul.
Gaming and interactive media
Game servers, matchmaking and rendering at edge PoPs close to players, with GPU capacity for streaming and content pipelines.
Fraud and network analytics
Models scoring SIP, SMS and signalling traffic in real time on GPUs beside the carrier core to detect IRSF, SIM-box, grey-route and smishing patterns.
Technical & commercial specifications.
Key parameters at a glance. Ask us for the full service description and SLA document.
| Edge locations | 40+ Dollu PoPs, partner colocation, on-premises appliances |
|---|---|
| Edge latency | < 10 ms metro round-trip; private backhaul to regions |
| Edge form factors | 1U single node to multi-rack; Kubernetes or VM hosts |
| GPU classes | NVIDIA L40S; H100 SXM with NVLink; fractional GPU for dev |
| GPU regions | Mumbai, Frankfurt, London, Ashburn, Singapore |
| Interconnect | Up to 400G between training nodes; NVMe scratch; parallel FS |
| Software | CUDA, PyTorch, TensorFlow, Triton, vLLM; managed Kubernetes |
| Consumption | On demand hourly, reserved 1–3 years, dedicated bare metal |
| Data controls | In-region only; customer-managed keys; no third-party telemetry |
| Support | 24×7 NOC; P1 response ≤ 15 min; named account manager |
How Cloud is priced.
Cloud and hosting are priced per resource - compute, storage, network egress, backup volume - with managed-service tiers on top. Migrations and DR projects are quoted as fixed scope.
We publish the model, not a public rate card - actual rates depend on destination, route class, volume and regulatory cost. See how every Dollu service is priced.
- Per resource: compute, storage, egress, backup volume
- Managed-service tiers per environment
- Reserved-capacity discounts on 12–36 month terms
- Migration and DR projects as fixed-scope quotes
- Billing
- Monthly recurring plus metered usage; itemised resources on every invoice
- Commitment
- On demand, or reserved for 12–36 months
Edge & GPU - your questions answered.
The questions customers and carriers ask us most often before they interconnect. If yours is not here, our team answers within one business day.
Still have a question?
Ask our solutions teamRelated services.
Services customers commonly combine with Edge & GPU Computing.
Bring your model to the edge.
Tell us your latency target, data locations and model size and we will propose PoPs, GPU classes and a network design within two business days.
