EndorsedendorsedCompleted

Self-hosted LLM & MLOps in private AWS for UK HealthTech

Services Covered:

AI and Machine Learning • Cloud Consulting and Services • DevOps
Project hero — replace with case study imagery
Industry
Industry

Software Development

Duration
Duration

2 months

Budget
Budget

10K - 49K

Start Date
Start Date

1 December 2025

Team Size
Team Size

1-5

Engagement Model
Engagement Model

Offshore

Project Capability Score (PCS)

gauge meter
82/100
Strong

Project Capability Score (PCS) estimates how capable Honeycomb Software is of successfully delivering a project like this, based on its past experience, track record, and ability to handle similar work. Learn more about PCS


Client

Sonomics

mapBarcelona, Spain

Project Summary

Honeycomb Software implemented a secure, scalable AI infrastructure for a custom software engineering firm (Sonomics) to enable a HealthTech client’s self-hosted LLM within a private AWS environment. They delivered MLOps architecture, CI/CD/GitOps pipelines, GPU autoscaling (scale-to-zero), observability, and thorough runbooks to meet UK data residency and compliance requirements. The work produced a production-ready, compliant LLM capability with clear handover documentation.

Key Challenges

  • Build a secure, scalable AI architecture the client can independently operate and expand
  • Run an LLM entirely inside the client's private AWS environment
  • Meet strict UK data residency and compliance requirements by keeping all AI processing and storage within the client's controlled infrastructure
  • Minimize GPU infrastructure costs while ensuring production readiness and observability

Project Deliverables

  • Kubernetes-based deployment of an open-source LLM in the client's private AWS environment
  • MLOps architecture design and environment setup
  • CI/CD pipelines and GitOps workflows for automated model deployments
  • Kubernetes/GPU node management and scale-to-zero autoscaling policies
  • Inference observability (metrics, logs, alerting) and operational runbooks
  • Structured handover and documentation

Project Solution

Acting as a specialist subcontractor, Honeycomb designed and implemented a Kubernetes-based deployment of an open-source LLM within the client’s private AWS environment. They provided MLOps architecture and environment setup, CI/CD pipelines and GitOps workflows for repeatable deployments, Kubernetes/GPU node management with scale-to-zero autoscaling to cut idle GPU costs, and an inference observability stack with metrics, logs, and alerting. The engagement included operational runbooks and a structured handover so the client’s team could operate and extend the platform independently.

Project Outcome

  • Compliant, production-ready self-hosted LLM running entirely within the client's UK AWS environment with no external API dependencies
  • Reduced idle GPU costs via scale-to-zero autoscaling
  • Full observability enabling confident production management of inference workloads
  • Clean handover and documentation allowing the client to operate and extend the platform independently
  • All deliverables were completed on time with clear communication and risk flagging

Platforms

  • CloudCloud

Tech Stack

  • Amazon Web Services (AWS)Amazon Web Services (AWS)
  • KubernetesKubernetes

Client Endorsement

Overall Review Rating

5star5 out of 5 stars

Timeliness

Cost Rating

Willing to Refer

Quality of Deliverables

“Honeycomb Software successfully delivered the project, enabling the client to have a compliant, production-ready self-hosted LLM capability. The team communicated clearly, took ownership of their deliverables, and flagged risks early. Their quality of handover documentation stood out.”

Iryna Stakhiv

BDM

Note: This endorsement is based on publicly available client feedback from external review sources.

More Projects by Honeycomb Software