Skip to content

Reading path

The articles in an order that makes sense.

The archive is sorted by date, and date is the wrong order for content that builds on itself. This is the order I would write in if I were teaching from scratch: each article assumes the previous one and sets up the next.

  1. 1Engineering5 minLatency, throughput and percentiles: the guide that settles the argumentYour average latency is lying to you. The three concepts that turn 'the system is slow' into a sentence with a number, an endpoint and a percentile.
  2. 2Engineering5 minCaching: the four traps nobody anticipatesAdding a cache is easy. The hard part is living with the four consequences it creates, and all four have known solutions.
  3. 3Engineering6 minIndexes, EXPLAIN and the slow queryHow an index works on the inside, why column order decides everything, and how to read an execution plan to know what to do.
  4. 4Engineering6 minSharding: a decision guideSharding is the most irreversible decision in a data system. This is the guide to deciding whether you need it and, if you do, how to pick the key.
  5. 5Engineering5 minCAP, PACELC and the consistency modelsCAP is the most quoted and most badly stated theorem in distributed computing. This article fixes the statement and shows the vocabulary you actually use day to day.
  6. 6Engineering5 minSaga and outbox: a transaction without a distributed transactionYou need to debit one account and credit another, and they live in different databases. The two patterns that solve it in practice, and what to avoid.
  7. 7Engineering6 minThe five resilience patterns that stop a cascading failureHow one slow secondary dependency takes down the whole system in ninety seconds, and the five patterns that prevent it.
  8. 8Engineering5 minSchema migration without downtimeWhy renaming a column breaks production, and the four step pattern that turns a migration into a Tuesday afternoon task.
  9. 9Engineering6 minObservability: metrics, logs and traces, and what each one answersThe three pillars are not interchangeable. Each answers a different question, and using the wrong one is why investigations drag on for hours.
  10. 10Applied AI6 minWhat an LLM really is: tokens, embeddings and attentionNo mysticism and no maths: what the model does, why it hallucinates, and how that changes your architecture decisions.
  11. 11Engineering5 minB-tree vs LSM-tree: inside your database's indexWhy Postgres is good at ranges and Cassandra is good at writes. It is not marketing: it is the tree each one picked, and the bill each choice comes with.
  12. 12Engineering6 minWhich data structure for which problemA decision guide organised by problem, not by structure. You have a need; which structure answers it, and at what cost.
  13. 13Engineering7 minConcurrency: race conditions, locks and deadlocksThe bugs that do not happen on your machine, do not happen in tests, and do happen in production. How to recognise them by reading code.
  14. 14Engineering5 minEvent loop or thread per request: how to chooseThe choice between the two models is not taste and not fashion. It is determined by the profile of your load, and there is a calculation that settles it.
  15. 15Engineering6 minWhy your service degrades over timeThe service starts fine and gets worse over days. Three causes explain almost every case, and all three have distinct signatures.
  16. 16Engineering6 minDiagnose the network in ten minutesA six command routine that turns "it must be the network" into "it is layer X, on hop Y".
  17. 17Engineering7 minThe vulnerabilities that actually show upMost breaches that make the news do not use sophisticated technique. They use one of these six flaws, and all six have a known, cheap fix.
  18. 18Engineering6 minOAuth2, OIDC and JWT demystifiedThe three most confused acronyms in authentication, what each one solves, and the mistakes that show up in almost every codebase.
  19. 19Engineering6 minTests worth what they costTests do not exist to prove the code is right. They exist so you can change it tomorrow without fear. That change of goal reorganises everything.
  20. 20Engineering7 minLegacy: characterise, seam and strangleThree techniques that let you safely change a system you did not write, do not understand, and cannot stop.
  21. 21Engineering7 minData and analytics: from OLTP to lakehouseTwo questions come up in every company: why did the report take down production, and why is the number on my dashboard different from yours. Both have the same root cause.
  22. 22Applied AI8 minAI in production: RAG, evals and prompt injectionAnybody can build a demo that impresses. What separates the demo from the product is three disciplines, and most teams have none of them.

Talk to me

Questions about the article? Message me on WhatsApp

No form and no mailing list. If you disagree with something I wrote, or want to tell me how you solved it, the conversation goes straight to me.

Open the chat