GenAI

AI-generated code: why “it works” does not mean “it is maintainable”

Coding agents quickly produce code that passes the tests. But code that works is not necessarily healthy code. What research on design smell detection teaches us, and how to keep technical debt under control.

October 12, 20267 min
Djamel Mesbah
CIFRE PhD candidate, Adservio
Key points
01 / Tests

Code that passes the tests can still be unhealthy

Tests check what the code does, not how it is designed. Code smells slip under that radar.

02 / Debt

The debt is paid later

These smells are associated with more changes, more defects and longer modification times.

03 / Detection

No model sees everything

Across 17 models tested in a PhD thesis carried out at Adservio, none scores above 0.30 on Feature Envy.

04 / Representation

The key is how code is represented

What the model is given to read matters more than its size or sophistication.

Code that passes the tests is not necessarily healthy code

Coding agents have changed the pace of software development. In a few minutes they write a function, a service, sometimes a whole module, that compiles and passes the tests. For a team under pressure, in banking, energy, healthcare or retail, it is tempting to consider the job done.

That is where the risk lies. A test checks a behaviour: for this input, that output. It says nothing about design: a method too long to understand, a class that centralises everything, logic duplicated in three places. The program works, but it becomes harder to understand, change and evolve.

These flaws have had a name since Martin Fowler's work: code smells, symptoms in the code that point to a deeper design problem. They matter all the more now that models generate working code very quickly, and that this code tends to accumulate poor design practices that are paid for, in the medium term, as technical debt.

This article draws on the research of Djamel Mesbah, carried out at Adservio as part of a CIFRE PhD with Université Paris-Saclay and ISEP, on the automatic detection of these smells with artificial intelligence. It has led to several scientific publications, listed at the end of the article.

Vibe Coding: Can Intuitive Programming Produce Production-Quality Software?
Related readVibe Coding: Can Intuitive Programming Produce Production-Quality Software?Vibe coding versus production quality: three experiments show what generative AI builds on its own, and why discipline and human oversight remain decisive.Read the article

Why design smells are costly

Code smells are not a matter of style. Many empirical studies associate them with lower maintainability and greater evolution effort.

More changes, more defects

Systems that concentrate smells such as God Class, Feature Envy or Long Method show more changes and defects, longer modification times and more complex change sets during maintenance. The components concerned accumulate technical debt: it becomes harder to understand the design intent, locate a change and propagate it safely.

A cognitive load that adds up

An isolated smell remains manageable for a developer. It is their accumulation that causes trouble: several co-existing smells clearly increase the workload and reduce comprehension accuracy. At the scale of a codebase continuously fed by agents, this accumulation is precisely what needs watching.

Four smells to know

01 / Method

Long Method

A method too long to be understood easily.

Best F1 score0.77
02 / Class

Blob, or God Class

A class that centralises too many responsibilities.

Best F1 score0.59
03 / Class

Data Class

A class that only holds data and accessors, with no real behaviour.

Best F1 score0.65
04 / Between classes

Feature Envy

A method more interested in another class's data than in its own.

Best F1 score0.28

Best F1 score reached across the 17 compared models, on a 0 to 1 scale. Source: Mesbah et al., Information and Software Technology, table 5.

Can detection be handed over to AI?

If AI produces these smells, can it also spot them? Research gives a nuanced answer, and that is what makes it useful for teams.

No model family dominates

Seventeen models from three broad families were compared on the same set of Java samples annotated by professional developers (the MLCQ dataset), under an identical protocol: classical learning on metrics, sequence models that read code as text, and neural networks on graphs built from the syntax tree. Graph-based models win on class-level smells (Blob, Data Class), sequence models on overly long methods. No family wins everywhere.

Fig. 01 · Smell detection

No model family dominates, and Feature Envy resists them all

Feature Envy
Long Method
Blob
Data Class
Classical models
0.23
0.69
0.55
0.54
Sequence models
0.28★
0.77★
0.48
0.57
Graphs (GNN)
0.12
0.72
0.59★
0.65★
Every family below 0.30

Best F1 score reached by each family on each smell (0 = no reliable detection, 1 = perfect detection). MLCQ dataset, 4,366 Java samples, mean of 10 runs. Source: Mesbah et al., Information and Software Technology, table 5.

Large language models: useful, but not enough

Queried directly through prompts, models such as GPT-4 and LLaMA reach limited performance on this task. Adding a few examples to the prompt clearly improves their answers, and GPT-4 remains the most reliable. But both struggle with Feature Envy and Blob.

One smell resists everything

Feature Envy remains an open challenge for every approach tested. The reason is telling: it is a relational smell. To see it, you need to know which classes call each other and with which types, information that today's code representations do not carry.

The lesson: it is not the model, it is the representation

The main takeaway fits in one sentence: what makes a smell detectable is not the sophistication of the model, but how the code it reads is represented.

Fig. 02 · Representing code

One piece of code, three ways to read it, one shared blind spot

Commande.java
class Commande {
  total() {
    client.remise
    client.pays
  }
}
View 01Metricssize, complexity, coupling, cohesion→ Classical models
× link not seen
View 02Token sequencecode read as text→ Sequence models
× link not seen
View 03Syntax graphthe structure of a class→ GNN
× link not seen
×Inter-class links · missing from all three viewsFeature Envy stays invisible

Diagram based on Djamel Mesbah's publications (AINA 2025, Information and Software Technology). The code and mini-visualisation values are illustrations.

More computing power does not change this. Distributing training reduces cost and computing time, but does not lift the detection ceiling: the limit lies in how code is represented.

For an IT department as for a development team, the lesson goes beyond research: a bigger model does not replace thinking about what it is given to see. That holds for smell detection, and it holds for coding agents themselves.

Measuring how clean generated code is

Reference benchmarks such as SWE-bench measure whether generated code works and solves the problem at hand. Another question remains largely open: is that code clean? Which design smells do models introduce into what they produce?

This is the direction explored by ongoing work at Adservio. The question becomes central as coding agents spread: producing code that passes the tests does not guarantee maintainable code.

How we integrate AI agents into our projects

At Adservio, these lessons shape how we integrate AI agents into our clients' software lifecycle, whatever their sector. Four principles structure our projects.

Fig. 03 · Integration pipeline

Check design at the moment debt is created

Illustration of this article's recommendations.

Measure design, not only behaviour

Tests remain essential, but they are no longer enough. Design smell detection enters the integration pipeline alongside the tests, so that technical debt becomes visible when it is created, not six months later.

Combine approaches rather than bet on one model

Since no model family dominates, robust tooling combines several approaches depending on the smell sought, rather than entrusting all detection to a single model, however powerful.

Keep humans on relational smells

Smells that require understanding relations between classes, such as Feature Envy, still escape tools. Design review by an experienced developer remains the safety net on these points.

Frame agents on proven patterns

A coding agent gives its best results on recurring tasks whose patterns have been validated and capitalised. We start with those cases, measure the gains, then widen progressively. The gain comes from framing the agents, not from setting them free.

Claude Code saved us 97% of the work, then it became a complete disaster
Related readClaude Code saved us 97% of the work, then it became a complete disasterAdservio's field report on Claude Code: Python support added to CodeConcise in minutes, a complete failure on JavaScript, and the lessons validated in 2026.Read the article

In short

Coding agents speed up software production, but speed says nothing about design quality. Measuring smells as soon as they appear and framing the agents is what turns a productivity gain into a lasting asset.

Sources

Mesbah D., El Madhoun N., Al Agha K., Zouaoui A., « A Survey on Code Smells Detection using Machine Learning Techniques », Information and Software Technology, Elsevier, accepted, 2026. https://www.sciencedirect.com/science/article/pii/S0950584926002314

Mesbah D., El Madhoun N., Al Agha K., Chalouati H., « Leveraging Prompt-Based Large Language Models for Code Smell Detection: A Comparative Study on the MLCQ Dataset », EIDWT, Springer, 2025. https://link.springer.com/chapter/10.1007/978-3-031-86149-9_42

Mesbah D., El Madhoun N., Al Agha K., Chalouati H., « Exploring NLP Techniques for Code Smell Detection: A Comparative Study », AINA, Springer, 2025. https://link.springer.com/chapter/10.1007/978-3-031-87769-8_9

Mesbah D., El Madhoun N., Al Agha K., Chalouati H., « Beyond the Code: Unraveling the Applicability of Graph Neural Networks in Smell Detection », NBiS, Springer, 2024. https://link.springer.com/chapter/10.1007/978-3-031-72325-4_15

Generative AICoding agentsTechnical debtSoftware quality

GET THIS ARTICLE

Download the full article as a PDF to read offline or share it.

SHARE THIS ARTICLE

On LinkedIn, X or by email, or just copy the link.

STAY POSTED

Get our next analyses and field notes straight to your inbox.

TALK TO AN EXPERT

Put these ideas into practice

Talk to our engineers about how this applies to your platform, your data and your teams.

By submitting this form, you agree to our privacy policy.

Frequently Asked Questions

A symptom in the code that signals a deeper design problem, as defined by Martin Fowler. The program works, but it becomes harder to understand, change and maintain. Examples: a method that is too long, a class that centralises too many responsibilities.

Code produced by LLMs often works, but it tends to accumulate poor design practices. Measuring exactly which ones is the subject of ongoing research at Adservio: passing the tests does not guarantee maintainable code.

Partly. Graph-based models detect class-level smells well, sequence models catch overly long methods, and GPT-4 does better with a few examples in the prompt. But some relational smells, such as Feature Envy, resist every approach tested.

By adding design smell detection to the integration pipeline, combining several detection tools, keeping human review on relational smells and framing agents on proven patterns.

Because code that is hard to evolve means a new feature that arrives later or costs more. Maintainability is a matter of time and budget, not only of engineering: the technical debt introduced today is paid with every change requested tomorrow.