Hi everyone!
In the first post we built a reference list of DDD's strategic and tactical patterns. Now for the question this community exists to explore: where can AI genuinely help with Domain-Driven Design, where might it hurt, and how would we know the difference?
Below is a map of research topics, grouped by the part of DDD they touch. Each one describes what the topic covers, what AI might contribute, and open questions worth investigating. None of these have settled answers, which is exactly what makes them interesting.
1. Domain discovery and knowledge crunching
1.1 Interview and workshop transcript analysis
Domain experts' knowledge usually surfaces in long, messy conversations. This topic explores how AI can process recorded interviews and workshop transcripts to extract candidate concepts, domain events, business rules, actors, and unresolved contradictions, with every extracted item linked back to the moment it was said. Open questions include how accurately AI separates genuine rules from casual remarks or one-off exceptions, whether it can capture tacit knowledge or only what's stated explicitly, and how the output should be presented so experts can review and correct it quickly.
1.2 Mining existing documents
Most organizations sit on requirements specs, regulations, policy manuals, contracts, support tickets, and email threads that all encode domain knowledge. This topic explores using AI, including retrieval over large document collections, to build a first-draft model from that material. The hard problems are handling conflicting sources, outdated documents, and the gap between documented process and how work is actually done, while keeping traceability so every concept can be justified by a source.
1.3 AI-assisted Event Storming
AI could play a role before, during, and after an Event Storming session. Before, it might generate a starter set of events from process descriptions. During, it could suggest missing events, spot inconsistent naming, flag hot spots, or act as a participant in remote sessions. Afterward, vision models could digitize a photo of the sticky-note wall into structured data, then cluster events into candidate aggregates and contexts. A central question is whether AI suggestions enrich the session or anchor participants too early and cut short the exploration that makes Event Storming valuable.
1.4 AI-assisted Domain Storytelling
This topic explores automatically turning stories narrated by domain experts into Domain Storytelling's pictographic notation of actors, work objects, and activities. AI could also generate variations of a story ("what happens if the payment fails?") to probe edge cases. Research questions include whether AI can reliably follow the notation's grammar and whether generated variations surface real scenarios or invented ones.
1.5 Gap detection and question generation
AI may be most useful not at answering questions but at asking them. Given a partial model, can AI identify undefined terms, missing lifecycle states, unhandled failure scenarios, and contradictions, then turn them into precise questions for domain experts? Expert time is usually the scarcest resource on a DDD project, so this could make every session more productive. Research directions include ranking questions by importance and measuring whether they actually lead to model improvements.
1.6 AI as a simulated domain expert
LLMs can role-play an insurance underwriter, a logistics dispatcher, or a nurse, drawing on general knowledge of an industry. This might help teams rehearse interviews, onboard developers, or build basic understanding of an unfamiliar field before meeting real experts. The obvious danger is confident, generic, or fabricated knowledge being mistaken for how a specific organization works. Research could explore when simulated experts are useful, how to clearly label their output as unverified, and how often they diverge from real experts.
2. Ubiquitous language
2.1 Glossary generation and maintenance
This topic explores AI building and continuously updating a glossary from conversations, documents, tickets, and code, with definitions, examples, discouraged synonyms, and the bounded context each term belongs to. The challenge is keeping humans as the authority over definitions while letting AI handle the tedious upkeep, and deciding how to treat terms that are still being debated.
2.2 Language drift detection
Over time, the words used in code, documentation, the UI, and everyday conversation drift apart. The code says Client while experts say "policyholder," or a ticket uses "shipment" for what the model calls a "consignment." Research could focus on AI tools that continuously compare terminology across these sources, flag divergence, and suggest renames or glossary updates. Measuring false-positive rates and whether teams act on the alerts would be key.
2.3 Ambiguity and polysemy detection
When one word carries different meanings in different conversations, it often signals a hidden bounded context boundary. This topic asks whether AI can detect these shifts in meaning, for instance noticing that "account" means a login in one discussion and a ledger in another, and propose them as candidate boundaries. It links language analysis directly to strategic design.
2.4 Naming assistance and language linting
AI could suggest class, method, and event names that align with the glossary, and review pull requests for names that bypass the ubiquitous language, such as technical jargon like DataManager or process() creeping into the domain layer. Open questions include how strict such linting should be and how to keep it from becoming noise that developers learn to ignore.
2.5 Multilingual ubiquitous language
Many teams discuss the domain in one natural language and write code in another, or work across several countries. This topic explores how AI can maintain precise, consistent translations of domain terms, and whether it can detect when a translation subtly changes a concept. Legal, financial, and regulatory terms are especially prone to this.
3. Strategic design
3.1 Subdomain classification
Given strategy documents, market information, and a description of the organization, can AI help classify subdomains as core, supporting, or generic and explain its reasoning? Since this is fundamentally a business judgment, the research question is less "can AI decide?" and more "can AI structure the discussion, challenge assumptions, and point out when a supposedly core area has actually become a commodity?"
3.2 Bounded context identification
Drawing context boundaries is one of the hardest and most consequential DDD decisions. This topic explores combining several signals to propose boundaries: semantic clustering of concepts from requirements, static analysis of code dependencies, co-change patterns from version control history, runtime call data, and team ownership. Existing approaches like Service Cutter use rule-based and clustering methods; the question is whether LLMs add useful semantic understanding on top. Proposals could be evaluated against the boundaries experienced architects would draw and against how well they hold up under real change.
3.3 Context map generation and maintenance
AI could reconstruct a context map from an existing system by reading code, API specifications, message schemas, and deployment configuration, then classifying relationships such as conformist, anticorruption layer, or open host service. The output could be written in a DSL like Context Mapper so it can be versioned and rendered. The longer-term challenge is keeping the map current automatically and detecting when an intended relationship, like an anticorruption layer, has quietly eroded.
3.4 Sociotechnical alignment
Conway's Law tells us that team structures and system boundaries shape each other. This topic explores AI analyzing commit authorship, code ownership files, org charts, and communication patterns to reveal misalignment, such as one context maintained by three teams or one team spread across five contexts. The findings could inform decisions about team structure as much as system design.
3.5 Distillation support
AI might assist with distillation by drafting a domain vision statement from strategy material, highlighting which parts of a codebase represent the core domain, identifying cohesive mechanisms worth extracting, and suggesting refactorings toward a segregated or abstract core. A key question is whether AI can tell genuinely differentiating logic apart from logic that's merely complex but generic.
4. Tactical design and code generation
4.1 Model-to-code generation
This topic covers generating entities, value objects, aggregates, repositories, and domain events from a model description or diagram, and comparing the results with human-written code on correctness, expressiveness, and faithfulness to the ubiquitous language. A known concern is that LLMs, trained on enormous amounts of CRUD-style code, tend to produce anemic models full of getters and setters rather than behavior-rich objects. Research could identify which prompts, examples, or constraints counteract this tendency.
4.2 Invariant discovery and aggregate design
Aggregates exist to protect invariants, so finding the right invariants is the real design work. This topic explores AI extracting invariants from business rules and conversations, proposing aggregate boundaries around them, and evaluating designs against guidelines like keeping aggregates small, referencing other aggregates by ID, and changing one aggregate per transaction. Open questions include whether AI can reason about concurrency and consistency trade-offs, and whether it recognizes when a rule can safely be eventually consistent.
4.3 Value object extraction
Primitive obsession, such as using strings for email addresses, decimals for money, or pairs of dates for ranges, is one of the most common weaknesses in domain models. AI could detect it in existing code, propose value objects with proper validation and behavior, and safely refactor the call sites. Because the problem is well defined and results are easy to verify, this is a promising early target for experiments.
4.4 Domain event design and evolution
This topic explores AI helping teams name events in the past tense using domain language, decide what data each event should carry, distinguish domain events from integration events, and manage schema changes and versioning over time. It also asks whether AI can detect weak events, such as commands disguised as events or vague events like OrderUpdated that carry no real domain meaning.
4.5 Business rule extraction into specifications
Rules written in plain language, like "a customer qualifies for premium shipping if they've placed more than five orders this year and have no overdue invoices," could be translated into composable specification objects. Research could measure the accuracy of this translation, explore ways to keep a traceable link between the written rule and its code, and test whether AI can flag rules that are ambiguous or contradict each other.
4.6 Anemic model detection and behavior relocation
AI could identify business logic sitting in application services, controllers, or utility classes that belongs inside entities and aggregates, then propose refactorings that move it where it belongs. An important part of this research is verifying that the resulting models are genuinely richer rather than simply reshuffled.
5. Refactoring and legacy modernization
5.1 Domain model recovery from legacy systems
Much critical domain knowledge lives only in legacy code, sometimes decades old, written in COBOL or PL/SQL, or buried in tangled monoliths whose original authors are long gone. This topic explores AI reverse-engineering the implicit domain model and business rules from such code and producing readable descriptions that current experts can validate. The core challenge is distinguishing intentional business rules from accidental behavior and historical workarounds.
5.2 Monolith decomposition
AI could propose seams for splitting a monolith into bounded contexts, plan step-by-step strangler fig migrations, and estimate the effort and risk of each extraction. Research could compare AI-proposed decompositions with those chosen by experienced teams and track how well each holds up after migration.
5.3 Anticorruption layer generation
Translation layers between models are tedious and error-prone to write by hand. This topic explores AI generating anticorruption layer code that maps between a legacy or external model and a clean domain model, along with tests that verify the translation. The interesting question is how well AI handles semantic mismatches rather than simple field mappings, such as when one system's "order" corresponds to two separate concepts in another.
5.4 Toward deeper insight
Evans describes "breakthroughs," moments when refactoring reveals a deeper and simpler model. This topic asks whether AI, by analyzing change history, recurring bugs, and awkward code, can point to places where the model is fighting the domain and suggest alternative ways of framing it. It's the most speculative topic on this list, but potentially one of the most valuable.
6. Testing, validation, and verification
6.1 Invariant-based test generation
AI could generate unit tests and property-based tests directly from aggregate invariants and specifications, checking that no sequence of operations can leave an aggregate in an invalid state. Research could measure how well these tests catch real defects compared with tests written by developers.
6.2 Acceptance scenarios in the ubiquitous language
This topic explores AI writing behavior-driven scenarios, for example in Gherkin, using the domain's vocabulary so that domain experts can read and validate them directly. That could close the loop between expert knowledge and executable tests. Useful measures include how often experts catch errors in generated scenarios and whether the scenarios reveal edge cases nobody had considered.
6.3 Model consistency and contradiction checking
AI could check a domain model against requirements, regulations, or known expert statements, flagging contradictions, missing cases, and rules implemented differently in different places. For critical invariants, this could be combined with formal verification techniques to provide stronger guarantees.
6.4 Architecture conformance and fitness functions
DDD designs erode as code changes: domain classes start depending on infrastructure, transactions span several aggregates, and contexts reach into each other's internals. This topic explores AI-assisted review that detects such violations, complementing rule-based tools like ArchUnit with the ability to recognize subtler design problems that are difficult to express as static rules.
7. Living documentation and knowledge sharing
7.1 Generated living documentation
AI could generate and continuously update documentation directly from the code and model, including context maps, aggregate diagrams, event catalogs, and glossaries that never go stale because they're rebuilt with every change. Research could explore how best to combine generated structure with human-written explanations of intent.
7.2 Domain onboarding and tutoring
This topic explores AI acting as a tutor that explains the domain to new team members using the organization's own glossary, code, and documentation. It could answer questions like "why can't an order be cancelled after dispatch?" by pointing to the relevant rule and code. Research could measure whether it shortens onboarding and improves newcomers' understanding of the domain, not just the codebase.
7.3 Decision capture
Design discussions happen in meetings, chats, and pull request comments, and the reasoning is often lost. AI could summarize these discussions into architecture decision records that capture the context, the options considered, and why one was chosen, so future developers understand why the model looks the way it does.
8. DDD as a structure for AI-assisted development
This category flips the question around: instead of asking how AI can help DDD, it asks how DDD can help AI.
8.1 Bounded contexts as boundaries for coding agents
AI coding agents struggle with large codebases partly because they can't hold everything in view at once. This topic explores whether scoping an agent's work to a single bounded context, with its own model, language, and clear interfaces, improves the quality and correctness of what it produces. Well-bounded contexts may turn out to be as valuable for AI agents as they are for human teams.
8.2 Ubiquitous language as context for AI
Does supplying the glossary, context map, and model documentation to an AI assistant measurably improve the accuracy of its code and answers? If so, DDD artifacts become a direct input to AI productivity, which could renew interest in practices that teams often let slide.
8.3 Aggregates as safe units of change
Aggregates define clear consistency boundaries and invariants. This topic explores whether they can serve as natural units of work for AI agents, limiting the scope of what an agent changes in one step and making those changes easier to review and verify.
8.4 Multi-agent systems organized by bounded context
This topic explores designing systems of cooperating AI agents in which each agent owns one bounded context and communicates through published languages and domain events rather than shared state. Context mapping patterns may apply directly, such as an anticorruption layer between agents with different models or an open host service one agent provides to many others.
8.5 Designing LLM-powered systems with DDD
When an application includes an LLM, where does it belong in the model? This topic explores treating the LLM as infrastructure behind a domain service interface, representing AI outputs as explicit domain concepts (for example, a Recommendation with a confidence level and review status), and keeping critical invariants enforced by deterministic domain code rather than by the model. It's about bringing DDD's discipline to the growing number of AI-powered products.
9. Evaluation and methodology
9.1 Benchmarks and reference domains
There are few public datasets of domain models with agreed-upon good answers, which makes comparing AI approaches difficult. This topic explores building benchmarks from well-known sample domains, like the cargo shipping example from the Blue Book or the sample applications from the Red Book, complete with expert-reviewed reference models, bounded contexts, and invariants.
9.2 Measuring model quality
How do you tell whether an AI-generated model is actually good? Research could combine several measures: expert review, coupling and cohesion metrics, invariant coverage, faithfulness to the ubiquitous language, and resilience when new requirements arrive. A model that looks tidy but breaks at the first real change request isn't a good model.
9.3 Human–AI collaboration dynamics
Much of DDD's value comes from the shared understanding a team builds while modeling together. This topic studies how AI affects that process. Does it speed up learning or let teams skip it? Do AI suggestions anchor discussions too early? Does the team end up understanding the domain better or worse? These questions call for observing real teams, not just comparing outputs.
9.4 Techniques and workflows
This topic compares practical approaches: prompting patterns, structured outputs, retrieval over domain documents, fine-tuning on an organization's code and terminology, and agentic workflows with tool use. The goal is to document which combinations work for which DDD tasks, with reproducible prompts and setups so others can build on them.
10. Risks and limitations
10.1 Hallucinated domain knowledge
LLMs can produce rules and concepts that sound entirely plausible but are wrong for a particular business. In a domain model, these errors become bugs built into the foundation of the system. This topic explores detection and mitigation, such as source traceability, confidence signals, and mandatory expert review at key points.
10.2 Loss of knowledge crunching
If AI can produce a model in minutes, teams may skip the conversations with domain experts that DDD depends on. The model would exist, but nobody would truly understand it. This topic explores how to use AI in ways that deepen collaboration with experts rather than replace it.
10.3 Pull toward the generic
LLMs learn from common patterns, so their suggestions naturally lean toward the typical. Yet the core domain is, by definition, what makes a business different. This topic investigates whether AI systematically flattens distinctive domain concepts into generic ones and how to preserve the nuances that matter most.
10.4 Over-engineering
AI assistants often apply sophisticated tactical patterns everywhere, even in generic or supporting subdomains where a simple CRUD approach would be the right choice. This topic asks whether AI can learn to judge when not to use DDD patterns, a skill experienced practitioners value highly.
10.5 Confidentiality of the core domain
The core domain represents a company's competitive advantage, and modeling it with AI can mean sharing sensitive business logic with external services. This topic explores the trade-offs between cloud and self-hosted models, data-handling policies, and techniques for getting AI assistance without exposing what makes the business unique.
Source: r/DDDWithAI · by /u/codeconjecture