Destello Tech · AI & ML Hardware
Run AI locally. Cut the cloud bill.
Professional GPU hardware for inference, LLM hosting, and GPU-accelerated compute. UK-stocked, next day dispatch, at a fraction of Nvidia enterprise pricing.
48GB
Max VRAM on a single board
UK
Stock, next day dispatch
3 SKUs
Arc Pro range in stock
Direct
Maxsun UK brand partner
Cloud GPU bills spiralling?
A single on-premise GPU can replace thousands in monthly cloud compute spend. The payback period on most AI workloads is under 6 months.
Data that can't leave your infrastructure?
GDPR, client confidentiality, IP protection. Local inference means your data never touches a third-party server. Full control, zero exposure.
Enterprise Nvidia pricing out of reach?
The Intel Arc Pro range delivers serious inference VRAM at a price point that makes on-premise deployment viable for startups and research labs, not just big enterprise.
The Intel Arc Pro Range
Purpose-built for AI inference
Purpose-built for AI inference
and professional compute.
All three SKUs are in UK stock now. Maxsun hardware, supplied direct through our brand partnership.
Arc Pro B60
24GB
GDDR6 VRAM
- Memory typeGDDR6
- InterfacePCIe 5.0 x16
- ArchitectureBattlemage Xe2
- XMX AI CoresYes
- Best for7B–13B models
Mid-range sweet spot
Arc Pro B70
32GB
GDDR6 VRAM
- Memory typeGDDR6
- InterfacePCIe 5.0 x16
- ArchitectureBattlemage Xe2
- XMX AI CoresYes
- Best for13B–34B models
Arc Pro B60 Dual
48GB
GDDR6 VRAM · Dual GPU board
- Memory typeGDDR6
- InterfacePCIe 5.0 x16
- ArchitectureDual Battlemage Xe2
- XMX AI CoresYes (×2)
- Best for70B+ models
What Can You Run?
VRAM estimates are approximate. FP16 = full precision. Q4 = 4-bit quantisation. Arc Pro XMX cores accelerate INT4/INT8 inference natively via OpenVINO and IPEX-LLM.
VRAM requirements at a glance.
The single most important spec for local LLM inference is VRAM. Here's what each GPU handles.
| Model | VRAM needed | B60 24GB | B70 32GB | B60 Dual 48GB | Notes |
|---|---|---|---|---|---|
| Llama 3.1 8B | ~16GB (FP16) | Full precision | Full precision | Full precision | Runs comfortably on all three |
| Mistral 7B | ~14GB (FP16) | Full precision | Full precision | Full precision | Excellent for RAG pipelines |
| Llama 3.1 70B | ~40GB (Q4) / 140GB (FP16) | Too large | Quantised only | Quantised (Q4) | Q4 quantisation gives ~90% quality |
| Mixtral 8×7B | ~26GB (Q4) | Quantised only | Full Q4 | Full Q4 | Strong for multi-task workloads |
| CodeLlama 34B | ~20GB (Q4) | Tight fit | Comfortable | Comfortable | Popular for code generation |
| Phi-3 Mini 3.8B | ~8GB (FP16) | Full precision | Full precision | Full precision | Ideal for edge / embedded use |
| Gemma 2 27B | ~16GB (Q4) | Comfortable | Full precision | Full precision | Strong reasoning performance |
Why Destello Tech
Direct pricing.
Direct pricing.
No runaround.
Maxsun UK Brand Partner
Direct trade access to the full Maxsun Arc Pro range. Pricing that most UK resellers cannot access — and no stock sourced from overseas.
UK Stock, Next Day Dispatch
Everything is held in the UK and ships from our Harlow workshop. No overseas lead times, no customs delays, no waiting weeks for a card to clear customs.
Direct to the Founder
You deal with Alex directly — no sales team, no ticket queue. Multi-unit orders, workstation builds, trade pricing — straight answer, fast turnaround.
Get in Touch
Tell us the workload.
Tell us the workload.
We'll spec the hardware.
Tell us what you're running, how many instances, and your budget. We'll come back with a GPU recommendation and pricing within 24 hours.
Want to talk it through? Book a quick call below.
Book a 15min Call
Alex Overend · Founder, Destello Tech
Email: alex@destellotech.com
Phone: +44 7756 296228
Harlow, Essex, UK