Skip to content
Cloud & Hosting

Edge & GPU Computing

Edge compute at Dollu PoPs and GPU instances for training and inference - so latency-critical applications and private AI for voice, contact-centre and analytics workloads run on your carrier's network, not across the internet.

Edge & GPU Computing - Dollu
  • 40+

    PoPs eligible for edge nodes

  • < 10 ms

    Edge round-trip in metro

  • H100 / L40S

    GPU classes available

Overview

What is Edge & GPU Computing?

Dollu Edge & GPU Computing places compute where distance matters. Edge nodes at Dollu PoPs and regional data centres run your containers and virtual machines within a few milliseconds of end users, branch sites, private 5G radios and the Dollu voice and messaging core, while GPU instances in our Mumbai, Frankfurt, London, Ashburn and Singapore regions provide the acceleration needed for model training, fine-tuning and high-throughput inference. Both are provisioned and operated with the same portal, API, Terraform provider and 24×7 NOC as the rest of Dollu Cloud.

Edge nodes are built for real-time work: local breakout for private 5G and SD-WAN traffic, media processing close to the caller, IoT ingestion and filtering, video analytics and application caching. Because they sit on the Dollu backbone, they can be stitched into your MPLS or SD-WAN VRF, receive traffic from Dollu DIA or Ethernet services and reach central regions over private paths. Nodes are deployed as managed Kubernetes clusters or VM hosts, from single-server footprints to multi-rack, in Dollu PoPs, colocation sites or on your premises.

GPU instances range from single L40S cards for inference and media workloads to multi-GPU H100 nodes with NVLink and high-bandwidth interconnect for distributed training. Instances come with pre-built images for CUDA, PyTorch, TensorFlow, Triton Inference Server and vLLM, NVMe scratch storage, high-throughput object storage for datasets and checkpoints, and optional managed Kubernetes with GPU scheduling. Capacity is available on demand, reserved or as dedicated bare-metal GPU servers.

Private AI is where the two meet. Speech recognition, speaker analytics, real-time translation, agent assist, call summarisation, fraud detection and conversational IVR all depend on low-latency access to live media and on keeping call content in-country and out of third-party APIs. Dollu runs these models on GPU capacity next to the voice core and at the edge, feeds them from Dollu CCaaS, SIP trunking and CPaaS media streams over private paths, and lets you bring your own models or use curated open-weight models under your control.

Why Dollu

Why choose Dollu for edge & gpu.

The advantages of buying from a carrier that owns its network, interconnects and operations - rather than a reseller.

  • Latency you can design around

    Metro round-trips under 10 milliseconds and deterministic private paths make real-time media processing, industrial control loops and interactive applications feasible without central-cloud delays.

  • AI on your data, in your jurisdiction

    Models run on GPUs in the region you choose, fed over private paths, so call recordings, transcripts and customer data never leave the country or pass through a third-party API.

  • Next to the voice core

    Speech and conversational AI workloads receive live media from Dollu SIP, CCaaS and CPaaS services with single-digit-millisecond latency, which makes real-time agent assist and translation practical.

  • GPU capacity without a capital project

    On-demand and reserved GPU instances, or dedicated bare-metal GPU servers, with pre-built software stacks - no data-hall build, power upgrade or hardware procurement cycle.

  • One network, one operator

    Edge nodes, GPUs, private 5G, SD-WAN and voice services from a single provider, monitored by one NOC, so the network path is engineered as part of the application rather than assumed.

Capabilities

Capabilities in detail.

Everything included with Edge & GPU Computing - the platform features, options and controls you get from day one.

  1. 01

    Edge node deployment options

    Managed edge clusters at Dollu PoPs, in partner colocation, or on-premises hardware we ship and operate; footprints from a single 1U server to multi-rack; Kubernetes or VM hosts.

  2. 02

    Network integration

    Edge nodes stitched into MPLS/SD-WAN VRFs, local breakout for private 5G user-plane traffic, anycast and GSLB steering to the nearest healthy node, private backhaul to Dollu regions.

  3. 03

    GPU instance families

    L40S instances for inference, media and graphics; H100 SXM nodes with NVLink and 400G interconnect for training; fractional GPU options for development; dedicated bare-metal GPU servers.

  4. 04

    AI software stack

    Images with CUDA, cuDNN, PyTorch, TensorFlow, Triton Inference Server, vLLM and NVIDIA NIM-compatible containers; managed Kubernetes with GPU operator, node pools and autoscaling.

  5. 05

    Data and storage

    Local NVMe scratch, high-throughput object storage for datasets and checkpoints, parallel file systems for multi-node training, and private paths from your data sources.

  6. 06

    Private AI for voice and contact centre

    Speech-to-text, real-time translation, agent assist, call summarisation, sentiment and compliance monitoring, and conversational IVR, deployed as services fed by Dollu voice media streams.

  7. 07

    Model lifecycle and MLOps

    Fine-tuning pipelines, model registry, A/B and canary serving, GPU utilisation monitoring and cost allocation per model or team; bring your own models or use curated open-weight models.

  8. 08

    Security and isolation

    Dedicated tenancy options, encrypted storage with customer-managed keys, private endpoints, no telemetry to third parties, and audit logs of every inference and training job.

How it works

How it works.

From first conversation to live traffic - a tracked, engineer-led onboarding with a named owner at every step.

  1. Step 01

    Define

    We capture latency targets, data locations, throughput and model requirements, and identify which PoPs, regions and GPU classes fit the workload.

  2. Step 02

    Provision

    Edge nodes are deployed at the selected PoPs or shipped to your sites; GPU instances or bare-metal servers are provisioned in-region with the software stack you selected.

  3. Step 03

    Connect

    Nodes and GPUs are attached to your VRF, private 5G core, voice media streams or data sources over private paths, with steering and failover configured.

  4. Step 04

    Deploy and tune

    Your team or ours deploys models and applications, validates latency and throughput against targets, and tunes batch sizes, quantisation and scheduling for cost.

  5. Step 05

    Operate

    24×7 NOC monitoring of nodes, GPUs and network paths, utilisation and cost reporting, and capacity planning as usage grows.

Use cases

Who uses this and why.

Typical deployments across carriers, enterprises, platforms and contact centres.

  • Contact centres and CCaaS

    Real-time transcription, agent assist, translation and post-call summarisation running on GPUs beside the voice core, keeping call content in-country and latency low enough to matter mid-call.

  • Conversational AI and voicebots

    Speech recognition, LLM reasoning and text-to-speech served from GPUs with sub-second turn-taking, connected directly to Dollu SIP trunks and CPaaS voice APIs.

  • Video analytics and smart sites

    Edge nodes at or near campuses, ports and factories process camera streams locally for safety, quality and security, sending only events and metadata upstream.

  • Manufacturing and private 5G

    Local breakout and edge compute for AGVs, machine vision and control systems on private 5G, with deterministic latency and no dependence on internet backhaul.

  • Gaming and interactive media

    Game servers, matchmaking and rendering at edge PoPs close to players, with GPU capacity for streaming and content pipelines.

  • Fraud and network analytics

    Models scoring SIP, SMS and signalling traffic in real time on GPUs beside the carrier core to detect IRSF, SIM-box, grey-route and smishing patterns.

Specifications

Technical & commercial specifications.

Key parameters at a glance. Ask us for the full service description and SLA document.

Edge locations40+ Dollu PoPs, partner colocation, on-premises appliances
Edge latency< 10 ms metro round-trip; private backhaul to regions
Edge form factors1U single node to multi-rack; Kubernetes or VM hosts
GPU classesNVIDIA L40S; H100 SXM with NVLink; fractional GPU for dev
GPU regionsMumbai, Frankfurt, London, Ashburn, Singapore
InterconnectUp to 400G between training nodes; NVMe scratch; parallel FS
SoftwareCUDA, PyTorch, TensorFlow, Triton, vLLM; managed Kubernetes
ConsumptionOn demand hourly, reserved 1–3 years, dedicated bare metal
Data controlsIn-region only; customer-managed keys; no third-party telemetry
Support24×7 NOC; P1 response ≤ 15 min; named account manager
Pricing model

How Cloud is priced.

Cloud and hosting are priced per resource - compute, storage, network egress, backup volume - with managed-service tiers on top. Migrations and DR projects are quoted as fixed scope.

We publish the model, not a public rate card - actual rates depend on destination, route class, volume and regulatory cost. See how every Dollu service is priced.

Monthly recurring · per resource
  • Per resource: compute, storage, egress, backup volume
  • Managed-service tiers per environment
  • Reserved-capacity discounts on 12–36 month terms
  • Migration and DR projects as fixed-scope quotes
Billing
Monthly recurring plus metered usage; itemised resources on every invoice
Commitment
On demand, or reserved for 12–36 months
Talk to salesActual rates within one business day.
FAQ

Edge & GPU - your questions answered.

The questions customers and carriers ask us most often before they interconnect. If yours is not here, our team answers within one business day.

Still have a question?

Ask our solutions team

At any of the 40+ Dollu PoPs with available space and power, in partner colocation facilities where you need a specific metro, or on your premises using hardware we ship, install and operate. All options connect over private paths to your VRF and to Dollu regions.

Let’s talk

Bring your model to the edge.

Tell us your latency target, data locations and model size and we will propose PoPs, GPU classes and a network design within two business days.

Abstract globe with connected network lines