Services
AI & LLM Infrastructure
Private deployment of language models and RAG pipelines. Production-grade vLLM setups with PagedAttention, Ollama clusters, ChromaDB integration, and low-bit model quantization (AWQ/GGUF) for running high-throughput inference without third-party cloud dependence.
Container Architecture & MLOps
Translating complex AI and application dependencies into clean, reproducible Docker Compose and Podman stacks. From isolated microservices to GPU pass-through containers, ensuring version-controlled and reliable deployments.
Linux Systems Engineering
Server setup, kernel tuning, security hardening, and ongoing administration. From initial bare-metal provisioning to performance tuning on Fedora, Debian, RHEL, and Ubuntu for mission-critical compute loads.
Infrastructure & Security Audit
A deep review of your current deployment — identifying bottlenecks, GPU utilization inefficiencies, and security flaws. Actionable guidance on container security, cost optimization, and strict system resilience.
Tools & Skills
AI & Inference
vLLM · Hugging Face · ONNX Runtime · CUDA · Ollama · Model Quantization (AWQ / GGUF) · ChromaDB · LlamaIndex · OpenWebUI · Python
Systems & Containers
Linux (Fedora, Debian, RHEL, Ubuntu) · Docker · Docker Compose · Podman (Rootless) · Systemd · Bash · Git · Kernel Tuning · Windows Server
Networking & Security
WireGuard · Tailscale · Nginx · Reverse Proxy · Cloudflare Tunnels · FIDO2 / Passkeys · DNS · Firewall Configuration
Observability & Monitoring
Prometheus · Grafana · Beszel · Netdata
Hardware & Edge
Raspberry Pi · ESP32 · Arduino · Circuit Design · Soldering · KiCad · FreeCAD
Business Software & Docs
Odoo ERP · Markdown · Inkscape · Microsoft Excel · Office 365
Dedicated AI workstation (Ryzen 9 + RTX 5070) and two Linux servers running 15+ self-hosted services. See the physical hardware and local model stack.
View the lab & hardware →Projects
LEMoE
↗AI middleware that semantically routes requests across multiple local expert models. Result: Up to 40% reduction in API inference costs and enables full execution on standard on-premise hardware.
Docker · Python · OllamaSRag
↗Containerized RAG pipeline connecting document store embeddings to a language model. Result: Achieved 100% data privacy and scalable architecture ready to process thousands of internal documents simultaneously.
ChromaDB · Docker ComposePRRPC
↗Linux daemon integrating hardware and software on Wayland. Drives a physical status monitor via a custom USB communication protocol. Result: Eliminated manual monitoring delays, providing real-time hardware status updates with sub-5ms latency.
Linux · Wayland · HardwareBlog
↗Technical writing in English and Spanish. Articles on Linux, infrastructure, automation, and hardware.
Writing · Linux · SystemsLab & Workstation
→The physical hardware running all of this. Dedicated AI workstation (Ryzen 9 + RTX 5070), two servers, 7.1 TB of storage, local GPU compute with no cloud dependency.
Ryzen 9 · RTX 5070 · HP EliteDesk · Raspberry Pi 4BLatest from the Blog
All articles ↗From Ollama to Production: Deploying vLLM with PagedAttention and Real-Time Metrics
A hands-on guide to deploying vLLM in production with Docker: optimize VRAM with PagedAttention, manage high concurrency, and monitor key telemetry metrics in real time.
Contact
If you need a reliable infrastructure partner — someone who understands your stack end to end — reach out.