You Pointed RAG at SharePoint. Nobody Owns What Goes Stale.

Release date:
September 6, 2026
Hero Vector
Abstract editorial illustration of tidy current documents with an amber check on indigo at left, and faded misaligned SharePoint pages with a crossed mark at right
Vector ImageVector ImageVector Image
Blog detail
Vector ImageVector ImageVector Image

The RAG pilot looked sharp. SharePoint was wired as the source of truth. Retrieval demos pulled the right policies. Leadership signed off. Six months later the assistant still answers fluently, and half the answers lean on pages nobody has touched since the restructure.

That is the part most RAG programmes underfund. Connecting SharePoint is an integration project. Keeping the corpus true is an operating model project. If you only finish the first one, you did not ship retrieval. You shipped a polite way to recycle stale intranet content.

Connection is not curation

Pointing a crawler at site collections feels like progress. Folders appear. Chunks land in the vector index. Citations look official. The missing question is quieter: who decides when a document is still fit to answer a customer, a staff member, or a regulator?

SharePoint is usually a collaboration surface, not a governed knowledge product. Drafts sit next to finals. Old versions hang around because deleting them feels risky. Project sites outlive the project. Regional teams keep parallel copies of the same procedure. RAG does not know the politics. It knows similarity.

In Australian enterprises this shows up as a familiar split. IT owns the connector and the index. The business owns the documents in theory. Knowledge management owns the taxonomy on a slide. Nobody owns the weekly decision that a page should be unpublished from the retrieval path.

If no named owner can take a document out of the answer path this week, you do not have a knowledge base. You have a museum with a search bar.

Where staleness actually lives

Stale RAG rarely fails with a red error. It fails with confident answers built from the wrong generation of truth:

  • Orphaned policy PDFs still indexed after the intranet page was updated
  • Team sites that were never meant to be enterprise guidance but score well on keyword overlap
  • Duplicate procedures across business units with slightly different rules
  • Archived projects whose status reports still look like current playbooks to an embedding model
  • HR and compliance packs that changed after a Fair Work or APRA update while old files remain searchable

Your evaluation set from go-live will not catch this for long. The questions stay the same. The documents move. Without freshness checks, the demos keep passing while production quietly drifts.

A pattern we keep seeing

Picture a national retailer that connected RAG to the policy SharePoint for store managers. Launch week was clean. Then a returns rule changed in one state. The intranet article was updated by the legal team. The PDF attachment in an old site library was not. The assistant kept citing the attachment because it chunked cleanly and ranked high. Store staff trusted the citation. The escalation queue filled with "but the AI said".

Or take a bank ops assistant pointed at procedures libraries. The connector ran nightly. Ownership of each library stayed with whoever created the folder years ago. When a payments workflow changed, the new SOP landed under a different site. The old SOP stayed live in the index for weeks because nobody owned a de-index request.

Neither case is a model failure. Both are corpus governance failures wearing a retrieval UI.

Treat the corpus like a product, not a crawl

Give RAG the same operational seriousness you give a customer channel.

  • Name owners per source collection. Not "SharePoint" as a blob. Named people for each library that can enter the answer path.
  • Define include rules. Final status, approved content types, and allowed sites only. Exclude drafts, personal drives, and archived hubs by default.
  • Track freshness SLAs. Every indexed source needs a review cadence and a last-verified date the ops team can see.
  • Prefer canonical pages. If both a wiki page and an attachment exist, retrieve the page and suppress the orphan file.
  • Make unpublish a first-class action. Business owners need a button, ticket, or workflow that removes content from the index without waiting for a full re-crawl debate.
  • Watch citation age. Alert when answers lean on documents past their review window or last edited beyond the SLA.

If your RAG runbook only covers chunk size, embedding model, and top-k, add a second chapter for corpus ownership. Make it as boring and mandatory as access reviews.

What to do before the next release

Before you expand the SharePoint footprint:

  1. List every site and library currently in the production index.
  2. Assign a business owner and a technical owner to each, with a freshness SLA in writing.
  3. Remove anything without an owner or without a review date this quarter.
  4. Add monitoring for citation age and duplicate near-matches across sites.
  5. Agree who can pull a document from the answer path the same day an error is found.

RAG against SharePoint can be excellent. It stays excellent only while someone is paid to keep the corpus honest. The connector was the easy part. Ownership of what goes stale is the product.

BG Image
Vector ImageVector ImageVector Image
We’re here to help
Vector ImageVector ImageVector Image

Ready to put AI to work in your business?

Talk to an AI expert about your goals.
Arrow Icon
Smart process automation
Arrow Icon
Direct access to our team. No bots.
Arrow Icon
We ask smart questions fast.

Book a Discovery Call

Your form has been submitted successfully. Thank you!
Please double-check your information and try again. If the issue continues, email us at info@trufyre.ai