Posts

Showing posts with the label Prompt Engineering

Token Cost Calculation: Why Your API Bill Is High and How to Fix It

Image
You ran a few tests, maybe built a small prototype, and then your OpenAI or Anthropic bill arrived — and it was way higher than you expected. Sound familiar? Understanding exactly how tokens are counted, how costs are calculated, and where the hidden waste lives is the single most impactful skill for anyone building with LLM APIs in 2026. This post walks you through the mechanics of tokenization, shows you how to measure usage precisely, and gives you battle-tested strategies to cut costs without sacrificing response quality. Table of Contents What Is a Token, Really? How Costs Are Calculated Counting Tokens Before You Send Where Tokens Hide: The Usual Suspects Strategies to Reduce Token Usage Prompt Compression Techniques Caching and Batching Monitoring and Budgeting in Production Putting It All Together 🔤 What Is a Token, Really? A token is not a word, and it is not a character — it sits somewhere in betw...

System, User, and Assistant Roles in the OpenAI Chat API Explained

Image
If you've ever called the ChatCompletion API and stared at the messages array wondering what system , user , and assistant actually do — you're not alone. These three roles are the backbone of every conversation you build with models like GPT-4o, and understanding them deeply unlocks everything from simple chatbots to production-grade AI assistants. By the end of this post, you'll know exactly what each role does, why it matters, and how to use them together in real code. Table of Contents What Is the Messages Array? The System Role The User Role The Assistant Role How the Three Roles Work Together Practical Code Examples Common Mistakes and How to Avoid Them Production Patterns Closing Summary 🗂️ What Is the Messages Array? When you call the ChatCompletion API, you don't send a single string — you send a list of message objects. Each object has two required fields: role and content . The...