Part book review, part argument about AI itself — and I’m genuinely not sure how much of it is which.
I have to start by saying that I’m not entirely sure how much of this essay—yes, a long text is incoming—is actually a review of the book and how much of it is me giving my opinions about AI.
I’ve made it a goal to read and review several books about AI this year. Mind you, I’m not an AI engineer or an AI specialist, and this is only the first book on my list. So, one could reasonably argue that I’m in no position to have strong opinions about the book or AI itself.
Fair enough.
But I’ve had my fair share of interactions with AI, and I’d like to think I’ve developed a few reasonably well-founded ideas along the way.
The first thing worth keeping in mind is that AI engineering is not a machine learning book. So, if you pick it up expecting to develop a fundamental understanding of machine learning from scratch, you might be a little disappointed.
You’ll certainly encounter plenty of ML concepts, and you’ll get a better idea of how things work, but quite a few sections may leave you wondering what something actually means at a deeper level or how exactly it is used in practice.
And honestly, that isn’t necessarily a bad thing.
As AI and foundation models become part of more and more products, every software engineer may eventually need at least a general understanding of AI without necessarily becoming a machine learning expert. ML engineering and AI engineering are, at least for now, related but distinct areas.
The book does a fantastic job of introducing concepts around how AI and foundation models work. You’ll learn about the rise of foundation models, sampling, model evaluation, model security, mitigating hallucinations, prompt engineering, and dataset engineering.
I found the sections on sampling and prompt engineering particularly useful. I think these will form a large part of how developers adapt foundation models to their applications.
And yes, I agree: “prompt engineer” probably isn’t a job title.
But prompt engineering is still a useful skill to have. It’s just nowhere near enough by itself to build production-ready applications.
The book also explains techniques such as RAG and fine-tuning reasonably well, although there isn’t quite enough information for you to go from reading the book to actually implementing them confidently in a real-world application.
When it comes to machine learning concepts, though, the book can become a little tedious.
I have mixed feelings about this part. Some sections introduce fascinating theoretical concepts, while others don’t really go very deep. At the same time, the book doesn’t necessarily follow a path that gradually builds a strong understanding of ML.
As someone who isn’t an ML engineer, I found some of the ML-related sections challenging to get through.
The author is very clear about one thing: quality assurance is essential for AI applications. Without it, the risks can easily outweigh the benefits.
She then goes through several techniques for evaluating models.
One of them is AI as a judge, where one model evaluates the output of another. The judge can be another LLM or a smaller, more specialized model.
The idea makes sense. After all, judging is easier than generating.
I mean, how many times have you seen people confidently judge someone else’s work without being able to do the job themselves?
Actually, the author also talks about getting creative when dealing with some of AI’s shortcomings. And every time she mentions that, I can’t shake the feeling that the suggestion of using AI to solve AI’s problems is coming right around the corner.
I’m obviously not saying that’s a bad idea in itself. It can be useful.
But I can’t help wondering whether this could eventually turn into a bit of a snowball effect.
Some of AI’s limitations
The book also discusses several limitations and problems with AI. A few of them really stood out to me.
1. Context is king.
And this has probably never been more true.
The importance of providing detailed and clear context to an LLM is huge. But, of course, you don’t get to have it that easy.
The longer the context becomes, the more likely the LLM is to focus on the wrong part of it.
There’s also something called position bias: LLMs tend to be better at following instructions placed at the beginning or end of a prompt than instructions buried somewhere in the middle.
One prompt-engineering technique is to repeat the original instruction after the user’s prompt.
And I thought only children had trouble remembering things unless they were the last thing they heard!
2. Have you ever felt that the more you study, the less you know?
Sometimes that feeling comes from being overwhelmed by the sheer amount of information available.
Sometimes, though, we simply forget things.
Fear not. You’re not alone.
Our LLM buddies are right there with us.
As an LLM learns more and more tasks, it becomes increasingly susceptible to something called catastrophic forgetting, where performance on previously learned tasks can deteriorate.
This might explain why I’m not very good at solving second-degree equations anymore.
Maybe I’m an LLM.
3. Okay, maybe I’m not an LLM.
There are already concerns that we may eventually run out of human-generated content with which to train LLMs.
Yes, LLMs are consuming publicly available information faster than new information is being produced.
This is where synthetic data—data generated by AI itself—is becoming increasingly important.
Synthetic data can be extremely useful when used carefully. The problem is that it can also produce a somewhat superficial form of learning.
A model trained on synthetic data generated by another model might learn how to provide a direct answer to a question without actually learning how or why that answer works.
And don’t expect it to admit that it doesn’t know.
Ask it to explain its answer and there’s a good chance it will simply hallucinate an explanation.
This reminded me of The Black Swan, where Nassim Nicholas Taleb discusses split-brain patients. When one hemisphere is instructed to perform an action and the other hemisphere is asked to explain why the action was performed, the patient may come up with a completely nonsensical explanation for it.
Anyway, there are studies correlating the use of synthetic data with model underperformance.
I bet that would make for one hell of a codebase!
It’s important to understand what AI brings to the table, as well as its implications and limitations.
The book does a good job of explaining this.
These limitations highlight many of the challenges we currently face with AI, which brings us to the broader discussion of AI in software development.
As many people already know, AI is probabilistic by nature.
Have you ever seen that meme saying, “Your chances are low, but never zero”?
That is basically AI’s motto.
Anything with a non-zero probability, no matter how wrong it is, can be generated by AI.
One issue I’ve personally experienced with AI is that it tends to reach conclusions without necessarily having factual consistency or sources to support what it is saying.
Just the other day, I was asking ChatGPT which tag I should use in an API I was consuming.
It confidently gave me an answer.
I then asked for the source.
It apologized and admitted that it didn’t actually have a specific source for the claim and that the answer was based on general knowledge.
And, of course, it turned out the answer was wrong.
To AI’s defense, people also make claims without sources all the time.
The difference is that I don’t generally see people apologizing and admitting that they made it up.
So… point for AI, I guess?
Still, I think this is a pretty major issue.
Before I challenged the answer, the model made no attempt to communicate that it was essentially guessing.
And that’s also where AI’s biases become easier to exploit.
I’ve found that AI can have a tendency toward confirmation bias—it often seems inclined to agree with the user’s premise.
Try asking it about the advantages of breaking a bone!
When an AI lacks factual consistency, it becomes surprisingly easy to steer its conclusion. And sometimes, it will happily fill the gaps with completely made-up information.
It’s almost like having a debate between two political parties, except the swerving is done by the AI.
Sadly, there are also wild LinkedIn coaches out there waiting to get you.
And, apparently, they have two particularly popular topics these days.
1. “Good software is delivered software.”
According to this school of thought, good design and architecture don’t really matter as long as the software gets delivered.
My fingers are already itching to dive deeper into this topic.
At the risk of going completely off-topic, I have to say that people who make this argument probably either don’t understand what good architecture actually means or have never had to deal with the consequences of a rushed product as a developer.
Probably both.
Rushed products almost always create far more cost in the long run than the revenue they generate in the short term.
2. “AI is really good at coding, and it will replace a significant number of developers soon.”
The coach gets bonus points if he checks all of these boxes:
- 2.1) He claims he built production-ready software in a week or less.
- 2.2) He did it completely by himself.
- 2.3) He doesn’t know how to code.
I’m not going to pretend I know exactly what the future of software development looks like.
But at the very least, someone making these kinds of claims and I probably have very different definitions of what “quality” means.
Good code requires thought and creativity.
If your first instinct when creating a new service or feature is to generate an MVC-like structure, then you and I probably have very different standards for what constitutes good software.
That doesn’t mean you shouldn’t use AI.
In my experience, AI can be genuinely useful for boilerplate, repetitive work, and some algorithms.
But, please, don’t ask it to design your domain model and blindly trust whatever it comes up with.
Of course, the response from the coach is usually something along the lines of:
You weren’t able to generate good code for a complex task? That’s a you problem. You need a better prompt.
Except you can literally paste a piece of code into the model, ask it to analyze it, and it will sometimes suggest an absurd change that clearly makes no sense and breaks the entire class.
Believe me, I’ve seen it happen.
Once, I watched one of these preachers arguing with a bunch of people about how developers are no longer needed.
He then proceeded to show off a “fairly complex” application—which was his own assessment of it.
It had a horrendous interface and a handful of buttons that basically did nothing more than calculate a few things.
I probably don’t need to mention that the person was a product guy with no software engineering background whatsoever.
What makes the whole situation worse is that people who talk like this are either deluded or trying to delude you.
Which one is which?
Just check their profile.
If they say they’re an AI specialist, they’re deluded.
If they say they own an AI product, they’re trying to delude you.
Good output requires good input.
Or, as data people love to say:
Garbage in, garbage out.
A large portion of an LLM’s training data consists of code samples.
So yes, it is absolutely possible to get an LLM to produce better code.
Again, though, I’d probably leave domain design out of that equation.
The important part is that you need to understand what good code actually looks like.
AI needs guidance to produce good output, and providing that guidance requires critical thinking.
And, at the risk of offending some people, I think critical thinking is generally a rather underrated—and unfortunately lacking—skill.
Sadly, God classes are also extremely common these days.
So you can imagine there is no shortage of them available as training material for AI.
On more than one occasion, I’ve seen AI demonstrate solid theoretical knowledge while being completely unable to apply that knowledge in practice.
One example was when I was discussing software anti-patterns.
ChatGPT was adamant that anemic domain models were problematic.
Fair enough.
So I asked it for some good examples.
It proceeded to generate anemic models.
When I tried to refine the code with it, it kept generating meaningless examples.
And that, to me, is one of the most interesting things about AI: knowing the definition of something and actually being able to apply that knowledge correctly are two very different things.
So, my conclusion?
We are doomed. AI will replace humans.
I mean, it already behaves like humans do:
- Full of biases.
- Judgmental.
- Spitting out claims without backing information.
- Losing focus the longer you talk to it.
- Forgetting stuff.
- And occasionally confidently explaining something it clearly doesn’t understand.
/sarcasm
Okay, here’s the actually useful conclusion.
This is a great book to start your AI engineering roadmap.
You might want to read some basic machine learning books first, especially if you don’t have an ML background, because that could make some of the ML sections less tedious.
But it’s not strictly necessary.
The book is dense, and there are times when it really manages to captivate you.
Overall, you’ll come away with enough knowledge to start thinking more seriously about how to refine AI applications. More importantly, you’ll have enough of a foundation to know what you need to learn next and where to look for more practical information on techniques such as RAG.
And just to be clear: my remarks about AI throughout this review are not meant to discourage people from using it.
Quite the opposite.
I genuinely believe AI has the potential to enable entirely new types of applications that we haven’t seen before.
AI can be incredibly useful.
But I also believe that AI is only as useful as the person using it.
A more skilled user, with stronger domain knowledge and better critical-thinking skills, is likely to get better results.
At the same time, I think it’s important for people to understand the limitations of LLMs.
It’s not as simple as taking an LLM and plugging it into your use case.
LLMs are not magic.
They are not a miracle.
And their output should absolutely not be treated as a source of truth.
Then again, neither should this review.
It’s just my opinion.
These are my own opinions, written after reading the book myself — not a summary, not a rating, and not sponsored by anyone.


