• The Future of AI Coding Isn’t Better Code—It’s Better Engineering

    The Real AI Coding Problem Nobody Is Talking About Why “hallucinations” are just the symptom — and the actual disease is something I call Engineering Drift.   Six months ago, I sat in a kickoff call for a Customer Support Platform. Nothing exotic. Tickets, chat, uploads, notifications, a billing plan, an admin dashboard. The kind…


  • Why a 7B Parameter Model Won’t Run Comfortably on a 14 GB GPU (And Why Most Engineers Get This Wrong)

    If you’ve recently started working with Large Language Models (LLMs), you’ve probably seen a calculation like this:   7 Billion Parameters × 2 Bytes (FP16) ≈ 14 GB   At first glance, it seems perfectly reasonable to conclude: “A GPU with 14 GB of VRAM should be enough.”   Unfortunately, that’s one of the most…


  • Introducing PersonaGraph: The Problem Every AI Power User Ignores

    moved to https://aitechpartner.com/introducing-personagraph-the-problem-every-ai-power-user-ignores/


  • RAG Evaluation Techniques — A Field Guide for Multi-Turn Voice Agents

    RAG Evaluation Techniques — A Field Guide for Multi-Turn Voice Agents   How to rigorously evaluate retrieval-augmented generation (RAG) in conversational voice agents — distinct from simple document Q&A interfaces. Written from real-world implementation experience with production voice systems.   All examples use the fictional company FleetPulse (fleet telematics provider) to illustrate concepts without referencing…


  • Why RAG Systems Sometimes Answer Questions Nobody Asked

    A Production Lesson Every AI Engineer Eventually Learns   One of the most surprising moments when deploying a Retrieval-Augmented Generation (RAG) system to production is watching users become frustrated by an AI that appears highly intelligent but somehow feels socially unaware.   The user says: Thank You The AI responds: According to the my knowledge…


  • The Hidden Context Window Problem in RAG Systems: A Real Production Incident with vLLM and Qwen3

    When Your 32K Context LLM Fails at 4K Tokens: A Production vLLM Troubleshooting Guide   One of the most common misconceptions in Generative AI systems is: “The model supports 32K context, so my application automatically supports 32K context.”   In production, that assumption can lead to unexpected failures.   Recently, we encountered a production issue…