Back to insights
AIEnterprise AIData GovernanceAI AgentsAgent MemoryRAG

AI Doesn’t Need More Memory. It Needs the Right Data

2026-06-30

By Igor Lima

Why useful AI agents must understand what is authoritative, what has changed and what is still true.

1. “Sorry, you’re right.”

You ask an AI assistant a question.

It gives you an answer.

You are not completely convinced, so you challenge it:

“That does not sound right. Are you sure?”

The assistant responds:

“Sorry, you’re right.”

It feels reassuring.

The model listened.

It accepted feedback.

It corrected itself.

Except sometimes you were not right.

Sometimes the AI had the correct answer, supported by the better source, and abandoned it simply because you sounded confident.

That might be harmless when asking about a film, a restaurant or a piece of trivia.

It becomes much more serious when the AI is connected to company information and helping people make decisions.

Imagine the approved product plan says:

The launch date is 21 July.

A user responds:

No, I am sure it is 14 July.

The assistant replies:

Sorry, you’re right. The launch date is 14 July.

The correct information was available.

The AI found it.

And then it gave it up.

Author's Perspective

The problem is not always that AI cannot find the truth. Sometimes it can find the truth but cannot hold on to it.

This is where the conversation about AI needs to move beyond which model is smartest.

Because an intelligent model operating over weak, outdated or poorly governed information can still make a very convincing mess.

2. Politeness can hide a real accuracy problem

The tendency for AI models to agree with a user, even when that agreement makes the answer worse, is usually called sycophancy.

It is not simply the model being friendly.

It is the model placing too much weight on what the user appears to believe.

Research note

Researchers tested nine language models across seven classification tasks. After being challenged with prompts such as “Are you sure?”, the models changed their answers 46% of the time on average, producing an average accuracy reduction of 17%.

Source: Are You Sure? Challenging LLMs Leads to Performance Drops in the FlipFlop Experiment

The worrying part is not that models sometimes reconsider an answer.

They should.

The worrying part is that the challenge itself can become more influential than the evidence.

A confident user can accidentally push the model away from the correct answer.

That problem becomes even more interesting when the AI has access to another source, such as an internal database.

Research note

A 2025 ACL study evaluated six language models when information supplied by the user conflicted with information retrieved from a database. The researchers found that user-provided information consistently had more influence. In some experimental conditions, models were up to three times more likely to be misled by incorrect user input than by incorrect database content.

Source: LLMs Trust Humans More, That's a Problem!

This creates an uncomfortable possibility.

We could invest heavily in connecting AI to trusted business systems, carefully retrieve the correct information, and still allow an unsupported comment from a user to override it.

The system has the right data.

But it does not understand which data deserves authority.

3. The real question is not only how intelligent the model is

Much of the AI conversation still centres on models.

Which one reasons better?

Which one performs best on benchmarks?

Which one has the largest context window?

Which one can remember the longest conversation?

These are reasonable questions.

But as AI enters real workflows, I think another question becomes more important:

Does the AI have the right version of reality?

A highly capable model can reason perfectly over the wrong information and still reach the wrong conclusion.

It can summarise an outdated policy beautifully.

It can write a confident customer response using an expired contract.

It can produce a detailed plan based on a decision that was reversed yesterday.

It can combine several individually correct facts into a final answer that is completely wrong.

This is why I increasingly see enterprise AI as a data problem as much as a model problem.

Author's Perspective

The model can only reason over the reality it is given—or believes it knows.

The model’s general training gives it the ability to write, interpret, reason and communicate.

But the organisation’s data tells it what is actually happening.

That includes information from:

  • CRM records
  • Emails
  • Customer tickets
  • Policies
  • Contracts
  • Project plans
  • Internal messages
  • Knowledge bases
  • Operational systems
  • Previous conversations
  • The agent’s own memory

These sources are not equally reliable.

They are not equally current.

And they do not always agree.

4. More information is not necessarily better information

A common response is to give the AI access to more.

More documents.

More systems.

More conversation history.

More memory.

More context.

The assumption is that if all the information is available, the model will work out the answer.

But access to more information can also mean access to more noise.

An old policy does not disappear when a new one is published.

A cancelled project plan can remain in SharePoint.

An early pricing spreadsheet can still sit in someone’s folder.

A temporary workaround can survive in a support article long after the original problem was fixed.

An informal Teams message can appear beside an approved decision.

The AI may retrieve all of them.

That does not mean it knows which one is right.

There is also evidence that simply placing more information into a model’s context does not guarantee that it will use it effectively.

Research note

Research published in Transactions of the Association for Computational Linguistics found that model performance could decline significantly depending on where relevant information appeared within a long context. Models often performed best when the information appeared near the beginning or end, and worse when it appeared in the middle.

Source: Lost in the Middle: How Language Models Use Long Contexts

A larger context window is useful.

But it is not the same as reliable understanding.

An AI can technically receive the correct information without giving it the right attention.

Author's Perspective

More context can increase what the model can see. It does not guarantee that the model knows what matters.

5. A response can be fully sourced and still be wrong

Consider a simple product launch.

The original plan says:

Launch date: 14 July
Location: London
Price: £49

The plan is later updated:

Launch date: 21 July
Location: Manchester
Price: £59

All six facts may continue to exist somewhere inside the company.

The original email is still available.

The first project plan remains indexed.

A marketing presentation still includes London.

Finance has approved the new price.

The venue team has confirmed Manchester.

The latest project record shows 21 July.

An AI searches for information about the launch and finds several relevant documents.

It responds:

The launch will take place in London on 21 July, with tickets priced at £59.

Every part of the response came from a real company source.

There may even be a citation beside every claim.

But that particular launch never existed.

The AI has created a new reality by combining different versions of the old and new plans.

Author's Perspective

A response can be fully sourced and still be wrong.

This is an important limitation of retrieval.

Retrieval is good at finding information connected to a question.

It is not automatically good at understanding organisational change.

It can identify documents about the launch.

It cannot always determine which document represents the current approved state.

Research note

Researchers studying outdated information in retrieval systems found that stale content could significantly reduce answer accuracy and mislead models even when current information was also available.

Source: HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation

That distinction matters.

The question is no longer only:

Did the AI retrieve a relevant source?

It also needs to ask:

Was it the right source for this decision, at this moment?

6. Organisational truth changes

Businesses do not operate from a permanent collection of fixed facts.

Reality changes every day.

Deadlines move.

Prices change.

Policies are replaced.

Customers change support tiers.

Employees move roles.

Risks are closed.

Incidents are resolved.

Approvals are withdrawn.

Temporary instructions expire.

A customer who was unhappy last month may now be satisfied.

A project described as “on track” on Monday may be at risk by Friday.

The information was not necessarily wrong when it was created.

It simply stopped being current.

This is particularly important for AI agents with memory.

We often talk about memory as though remembering more must make an AI more useful.

But a system that remembers everything without understanding change can become less useful over time.

Old instructions remain active.

Temporary preferences become permanent assumptions.

Incorrect interpretations are carried into future tasks.

Previous summaries become detached from their original evidence.

The AI remembers what happened.

But not necessarily what still applies.

Research note

The 2026 STALE benchmark tested AI agents across 400 expert-validated scenarios and 1,200 questions involving changing information. Even the strongest evaluated model achieved only 55.2% overall accuracy. Models frequently retrieved updated evidence but failed to apply it, particularly when the user’s question contained an outdated assumption.

Source: STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?

That last point connects directly to the “sorry, you’re right” problem.

Suppose the AI knows that the launch moved from London to Manchester.

The user later asks:

What time should the team arrive at the London venue?

A strong system should recognise that the question contains an outdated premise.

It should not confidently provide directions to London.

It should say:

The latest approved plan shows the venue has changed to Manchester. Would you like the arrival information for the new venue?

That requires more than memory retrieval.

It requires the AI to understand that reality changed and that the earlier information should no longer guide the answer.

7. The right data has qualities that “more data” does not

When I say AI needs the right data, I do not mean a perfectly clean database with every possible field completed.

I mean that the system needs enough information to judge what deserves to influence the answer.

The right data should be:

Authoritative

Is this the approved source, or an informal conversation?

An email can be authoritative evidence that somebody said something.

It may not be authoritative evidence that what they said was correct.

Current

Is this still true?

A policy may have been completely accurate when written and still be wrong for today’s decision.

Relevant

Does it apply to this customer, country, product, contract or point in time?

A correct policy for one market can be the wrong policy for another.

Permissioned

Is the user—and the agent acting for that user—allowed to access and use this information?

Being able to retrieve something does not automatically mean it should be included in an answer.

Traceable

Can the organisation see where the information came from?

Could somebody inspect the source if the decision was challenged?

Revisable

Can the information be updated, retired or replaced when reality changes?

A memory system that can only add new facts will eventually become full of competing truths.

Research note

NIST’s AI Risk Management Framework emphasises transparency, accountability, human oversight and data provenance. It notes that maintaining provenance can support both transparency and accountability, while documentation should continue across the AI lifecycle as systems and their operating contexts evolve.

Source: NIST Artificial Intelligence Risk Management Framework

These qualities turn information into something the AI can use responsibly.

Without them, the model sees text.

With them, it begins to see context, ownership, time and authority.

8. Human feedback is important—but it is not automatically ground truth

I strongly believe in keeping humans involved in important AI workflows.

But human-in-the-loop should not mean:

The human is always right.

Humans misremember.

We use old information.

We confuse similar projects.

We make assumptions.

We sometimes push for the answer we want rather than the answer the evidence supports.

A mature AI system should therefore distinguish between several different things.

A challenge

Are you sure?

This should trigger a check.

It should not automatically trigger an apology and reversal.

New evidence

Here is the approved plan showing that the date changed.

This gives the system something new to evaluate.

An authorised update

I am the product owner and this has now been formally changed.

This may justify updating the current state, depending on the organisation’s process.

An unsupported assertion

I know the old date is correct.

Confidence alone should not override a stronger source.

This does not mean designing AI that argues with people.

It means designing AI that is open to correction without becoming obedient to every confident statement.

Research note

Anthropic found sycophantic behaviour across five leading AI assistants and four different text-generation tasks. Its analysis also found that responses matching a user’s views were more likely to be preferred, and that both humans and preference models sometimes preferred convincingly written agreement over correct answers.

Source: Towards Understanding Sycophancy in Language Models

This may partly explain why the behaviour emerges.

Agreement often feels better.

A response that validates us can appear more helpful than one that questions our assumption.

But helpfulness and agreement are not the same thing.

Author's Perspective

A good AI should be easy to correct, but difficult to push away from stronger evidence.

9. This is not only a theoretical problem

In April 2025, OpenAI rolled back an update to GPT-4o after the model became noticeably too agreeable.

Research note

OpenAI said the removed update had made GPT-4o overly flattering and agreeable. The company explained that it had placed too much emphasis on short-term user feedback, resulting in behaviour that felt supportive but could be disingenuous.

Source: OpenAI: Sycophancy in GPT-4o

That is worth paying attention to.

These systems are partly shaped by our reactions to them.

We tend to reward answers that feel useful, supportive and aligned with us.

But an AI that optimises too heavily for the immediate positive reaction may become less willing to challenge a false premise.

In a personal conversation, that can create misplaced confidence.

In an enterprise workflow, it can affect customer communication, operational decisions, compliance, financial information or risk management.

The model should not be difficult for the sake of being difficult.

But sometimes the most helpful answer is:

The approved information does not support that.

Or:

I found two conflicting versions and cannot safely determine which is current.

Or:

The source you mentioned is older than the current policy.

Or simply:

I may be missing an update, but I should not change this answer without stronger evidence.

That is not a failure of the AI.

That is the behaviour of a system designed to know its limits.

10. Better AI requires better organisational clarity

It is tempting to see this as a purely technical challenge.

Improve retrieval.

Add timestamps.

Use a better database.

Create more sophisticated memory.

All of that can help.

But technology cannot invent clarity that the organisation does not have.

If nobody knows which system is authoritative, the AI will not know either.

If policies are copied across several repositories, the AI inherits that duplication.

If important decisions happen only in informal chat, the AI cannot reliably distinguish discussion from approval.

If nobody owns the process for retiring obsolete information, old facts will remain available indefinitely.

If access permissions are inconsistent, the agent may see information it should not use.

AI often exposes problems that already exist in the organisation.

It does not create the messy document library.

It makes the consequences of that mess more visible.

Author's Perspective

AI does not make organisational information messy. It makes the consequences of that mess harder to ignore.

This may be one reason why AI looks so impressive in a controlled demonstration and becomes harder in production.

The demonstration has five clean documents.

The production environment has:

  • Three versions of the policy
  • Two systems claiming to be the source of truth
  • A decision hidden in an email thread
  • A temporary workaround that became permanent
  • Missing ownership
  • Outdated permissions
  • Conflicting customer records
  • A user confidently insisting that the old answer is correct

The model may be functioning exactly as designed.

The information environment around it is not.

11. What organisations should be asking

Before connecting an AI agent to more systems, I think organisations need to ask a few less glamorous questions.

Which source is authoritative for each important decision?

How will the AI know when information has changed?

What happens to the previous version?

Who is allowed to correct a stored fact?

Does a user’s comment become permanent memory?

Can the system explain why one source was trusted over another?

What happens when an approved source and a human instruction conflict?

Can the AI say that it does not have enough evidence?

Can people inspect and challenge the decision without destroying the audit trail?

These questions may not attract the same attention as a new model release.

But they will determine whether the AI can be trusted inside a real workflow.

A reliable system needs to be able to separate:

  • What was said
  • What was inferred
  • What was approved
  • What was temporary
  • What has changed
  • What is true now

That is the foundation underneath the experience.

12. My hypothesis

The next major improvement in enterprise AI will not come only from larger models, longer context windows or more persistent memory.

It will come from better data discipline around the model.

That means:

  • Identifying authoritative sources
  • Tracking when facts become valid
  • Preserving provenance
  • Separating evidence from inference
  • Recognising contradictions
  • Updating memory when reality changes
  • Retiring information that should no longer guide decisions
  • Resisting unsupported assumptions, including those from confident users
  • Keeping an audit trail of what changed and why

The goal is not an AI that remembers everything.

It is an AI that can answer three simple questions:

What did we believe before?

What changed?

What is true now?

Conclusion

“Sorry, you’re right” sounds harmless.

Sometimes it is exactly what an AI should say.

But sometimes those four words reveal a deeper weakness.

The model is not following the evidence.

It is following us.

As AI becomes more connected to company systems, more personalised and more capable of acting, this distinction matters.

A larger context window can hold more outdated information.

A better retrieval system can find more conflicting records.

A longer memory can preserve more assumptions that no longer apply.

A more agreeable model can be persuaded to abandon the correct answer.

Good AI therefore depends on more than intelligence.

It depends on whether the system understands authority, relevance, permission, provenance and time.

It must remain open to correction without treating confidence as evidence.

It must remember previous decisions without becoming trapped by them.

And when the data changes, the AI needs to change with it.

Author's Perspective

A good AI agent does not need to remember everything. It needs to know which information still deserves to be remembered.

Author's note: The views and opinions expressed in this article are my own and are based on a combination of published research and personal observations. References are provided where applicable.

Want to explore a similar workflow?

Try the AI demos or explore how practical AI workflows can support operations, knowledge retrieval and human-in-the-loop decision making.