Code that passes the tests is not necessarily healthy code
Coding agents have changed the pace of software development. In a few minutes they write a function, a service, sometimes a whole module, that compiles and passes the tests. For a team under pressure, in banking, energy, healthcare or retail, it is tempting to consider the job done.
That is where the risk lies. A test checks a behaviour: for this input, that output. It says nothing about design: a method too long to understand, a class that centralises everything, logic duplicated in three places. The program works, but it becomes harder to understand, change and evolve.
These flaws have had a name since Martin Fowler's work: code smells, symptoms in the code that point to a deeper design problem. They matter all the more now that models generate working code very quickly, and that this code tends to accumulate poor design practices that are paid for, in the medium term, as technical debt.
This article draws on the research of Djamel Mesbah, carried out at Adservio as part of a CIFRE PhD with Université Paris-Saclay and ISEP, on the automatic detection of these smells with artificial intelligence. It has led to several scientific publications, listed at the end of the article.

Why design smells are costly
Code smells are not a matter of style. Many empirical studies associate them with lower maintainability and greater evolution effort.
More changes, more defects
Systems that concentrate smells such as God Class, Feature Envy or Long Method show more changes and defects, longer modification times and more complex change sets during maintenance. The components concerned accumulate technical debt: it becomes harder to understand the design intent, locate a change and propagate it safely.
A cognitive load that adds up
An isolated smell remains manageable for a developer. It is their accumulation that causes trouble: several co-existing smells clearly increase the workload and reduce comprehension accuracy. At the scale of a codebase continuously fed by agents, this accumulation is precisely what needs watching.
Four smells to know
Long Method
A method too long to be understood easily.
Blob, or God Class
A class that centralises too many responsibilities.
Data Class
A class that only holds data and accessors, with no real behaviour.
Feature Envy
A method more interested in another class's data than in its own.
Best F1 score reached across the 17 compared models, on a 0 to 1 scale. Source: Mesbah et al., Information and Software Technology, table 5.
Can detection be handed over to AI?
If AI produces these smells, can it also spot them? Research gives a nuanced answer, and that is what makes it useful for teams.
No model family dominates
Seventeen models from three broad families were compared on the same set of Java samples annotated by professional developers (the MLCQ dataset), under an identical protocol: classical learning on metrics, sequence models that read code as text, and neural networks on graphs built from the syntax tree. Graph-based models win on class-level smells (Blob, Data Class), sequence models on overly long methods. No family wins everywhere.
No model family dominates, and Feature Envy resists them all
Best F1 score reached by each family on each smell (0 = no reliable detection, 1 = perfect detection). MLCQ dataset, 4,366 Java samples, mean of 10 runs. Source: Mesbah et al., Information and Software Technology, table 5.
Large language models: useful, but not enough
Queried directly through prompts, models such as GPT-4 and LLaMA reach limited performance on this task. Adding a few examples to the prompt clearly improves their answers, and GPT-4 remains the most reliable. But both struggle with Feature Envy and Blob.
One smell resists everything
Feature Envy remains an open challenge for every approach tested. The reason is telling: it is a relational smell. To see it, you need to know which classes call each other and with which types, information that today's code representations do not carry.
The lesson: it is not the model, it is the representation
The main takeaway fits in one sentence: what makes a smell detectable is not the sophistication of the model, but how the code it reads is represented.
One piece of code, three ways to read it, one shared blind spot
Diagram based on Djamel Mesbah's publications (AINA 2025, Information and Software Technology). The code and mini-visualisation values are illustrations.
More computing power does not change this. Distributing training reduces cost and computing time, but does not lift the detection ceiling: the limit lies in how code is represented.
For an IT department as for a development team, the lesson goes beyond research: a bigger model does not replace thinking about what it is given to see. That holds for smell detection, and it holds for coding agents themselves.
Measuring how clean generated code is
Reference benchmarks such as SWE-bench measure whether generated code works and solves the problem at hand. Another question remains largely open: is that code clean? Which design smells do models introduce into what they produce?
This is the direction explored by ongoing work at Adservio. The question becomes central as coding agents spread: producing code that passes the tests does not guarantee maintainable code.
How we integrate AI agents into our projects
At Adservio, these lessons shape how we integrate AI agents into our clients' software lifecycle, whatever their sector. Four principles structure our projects.
Check design at the moment debt is created
Illustration of this article's recommendations.
Measure design, not only behaviour
Tests remain essential, but they are no longer enough. Design smell detection enters the integration pipeline alongside the tests, so that technical debt becomes visible when it is created, not six months later.
Combine approaches rather than bet on one model
Since no model family dominates, robust tooling combines several approaches depending on the smell sought, rather than entrusting all detection to a single model, however powerful.
Keep humans on relational smells
Smells that require understanding relations between classes, such as Feature Envy, still escape tools. Design review by an experienced developer remains the safety net on these points.
Frame agents on proven patterns
A coding agent gives its best results on recurring tasks whose patterns have been validated and capitalised. We start with those cases, measure the gains, then widen progressively. The gain comes from framing the agents, not from setting them free.

In short
Coding agents speed up software production, but speed says nothing about design quality. Measuring smells as soon as they appear and framing the agents is what turns a productivity gain into a lasting asset.
Sources
Mesbah D., El Madhoun N., Al Agha K., Zouaoui A., « A Survey on Code Smells Detection using Machine Learning Techniques », Information and Software Technology, Elsevier, accepted, 2026. https://www.sciencedirect.com/science/article/pii/S0950584926002314
Mesbah D., El Madhoun N., Al Agha K., Chalouati H., « Leveraging Prompt-Based Large Language Models for Code Smell Detection: A Comparative Study on the MLCQ Dataset », EIDWT, Springer, 2025. https://link.springer.com/chapter/10.1007/978-3-031-86149-9_42
Mesbah D., El Madhoun N., Al Agha K., Chalouati H., « Exploring NLP Techniques for Code Smell Detection: A Comparative Study », AINA, Springer, 2025. https://link.springer.com/chapter/10.1007/978-3-031-87769-8_9
Mesbah D., El Madhoun N., Al Agha K., Chalouati H., « Beyond the Code: Unraveling the Applicability of Graph Neural Networks in Smell Detection », NBiS, Springer, 2024. https://link.springer.com/chapter/10.1007/978-3-031-72325-4_15
STAY POSTED
Get our next analyses and field notes straight to your inbox.



