AI, LLMs, and applied ML.
Senior-engineer field notes on AI, LLMs, agents, and applied machine learning by Ercan Ermis.
All posts
One Door to the Models, Part 2: Terraform, Bicep, or ARM
Bicep has no state file and cannot touch Entra ID. Part 2 picks the IaC layer for the gateway, and finds the model-version default that upgrades production.
One Door to the Models, Part 1: The Case for a Central LLM Gateway
A company runs five GenAI apps and cannot say what any of them cost. Part 1: the scenario, build versus buy on Azure, and the gateway this series builds.
Agents on Call, Part 8. Production: Observability, Evals, and the Day It Lies
OTEL traces into CloudWatch, Bedrock invocation logging to S3, an eval harness with a golden incident set, and the day the triage agent lied with confidence.
Kiro After the Hype: What AI IDEs Actually Changed
Eight months after Kiro went GA, the spec survived and the IDE mostly did not. What actually changed is where the review happens, not who writes the code.
Why Your AI Pilot Died in Procurement
Most AI pilots do not fail on accuracy. They stall in security review and data processing agreements, because nobody scheduled the twelve weeks that follow.
Per-App Bedrock Cost Tracking with Inference Profiles
Application inference profiles put cost allocation tags on Bedrock calls, turning one shared bill into per-team lines. The tag design is the hard part.
Agents on Call, Part 7. Sizing: Token Math Nobody Does Upfront
The sizing math nobody does upfront: tokens per incident, quota ceilings, when provisioned throughput breaks even, and the platform's monthly bill.
EU AI Act, August 2: The Deadline That Didn't Move
The Digital Omnibus deferred the high-risk deadlines to 2027 and 2028. Article 50 transparency still lands on 2 August 2026, and that is the one most teams hit.
SageMaker vs Bedrock: An Org Decision, Not a Technical One
The SageMaker or Bedrock question is really about whether your org has a team that owns models. Pick for the team topology you have, not the one on the slide.
Prompt Injection via Your Own Docs: The RAG Attack Surface
Your knowledge base is untrusted input. Retrieval hands attacker-authored text to the model, so the control that pays off is scanning at ingestion time.
More from Ercan
Two more sites, same author, different ground.
Cloud, AWS, EKS, Terraform, platform engineering.
Field notes from production systems. EKS, IAM, Terraform at organization scale, observability, cost optimization.
Visit ercan.cloud →The hub. About, consulting, contact.
Personal hub for both writing tracks. Who I am, how the consulting works, how to reach me.
Visit ercanermis.com →