AI Visibility and MediaWiki

From mw.mh370.wiki
Jump to navigation Jump to search

AI Visibility and MediaWiki

If you are planning a new MediaWiki site on the Internet, consider how to gain AI Visibility before you start.

If, like the mh370wiki.net website, the site already exists, a strategy to gain AI Visibility will be more involved.

Issues with a mature MediaWiki website

I upgraded the site to Semantic MediaWiki thinking all I had to do to modernise it was to add semantic properties... unfortunately, there is more to it!

To use mh370wiki.net as an example, the site has been developed since 2015. Unlike most MediaWiki sites, all content uses HTML and CSS; there is no native wikitext. During the past decade HTML5 has been updated but there is no version 6. Various tags have been 'deprecated' and new tags have been introduced. Browsers may be backward-compatible, but the reality is - a mature site needs an audit and some update. Fortunately, the mh370wiki.net has minimal in-line CSS. However, all of the content quoted from official sources is enclosed in a custom quote class because I didn't like the way blockquote looked. To correctly inform an AI agent that will need to change.

Some Surprises

HTML has been updated and now includes many semantic tags which identify parts of a webpage. For example, (enclosed in <>) article, header, footer, summary, aside etc. Operators of normal websites are encouraged to incorporate semantic html elements to assist AI agents understand the structure of webpages.

The surprise for MediaWiki developers is:-

  1. We cannot easily include these tags, and
  2. we don't need to.

With $wgRawHtml = true; in LocalSettings.php these semantic tags become part of the visible page. The only way to hide them is to enclose the entire page in html tags but the consequence is that transclusion does not work inside html tags, and therefore templates won't work either.

However, MediaWiki has been updated too, and a content page is apparently presented for us as an article. This applies to pages created with wikitext as well as those using html.

And another surprise - MediaWiki may already inject some json-ld script in the webpage head section. Here are some excerpts from the script on this page:-

<script type="application/ld+json">{"@context":"http:\/\/schema.org","@type":"Article",
"name":"AI Visibility - mw.mh370.wiki",
"dateModified":"2026-09-26T12:40:36Z",
"datePublished":"2026-09-26T12:40:36Z",
"author":{"@type":"Organization","name":"mw.mh370.wiki",
publisher":{"@type":"Organization","name":"mw.mh370.wiki",
"potentialAction":{"@type":"SearchAction","target":"https:\/\/mw.mh370.wiki\/w\/index.php?title=Special:Search&search={search_term}","query-input":"required name=search_term"}}</script>
</head>

This is a bit too generic and a goal would be to provide more valuable information about the website, organisation, author, publisher etc.

Note: An Internet search for more information about how this script is in my page head or where the data is stored was not helpful. One answer was quite authoritative: "In MediaWiki, the <script type="application/ld+json"> tag does not exist in core by default. Standard MediaWiki installations do not output JSON-LD structured data."

Headings and Semantic HTML

It is recommended that webpages use headings h1, h2 etc in order and with no level missing.

In MediaWiki the page name is captured within h1 tags so the highest level heading on any content page should be h2. Sub-headings should therefore be h3, h4 and h5 etc.

The Way Forward

Conceptually, it may seem difficult to view a website about an event, like mh370wiki.net as data. It is easier to accept that a website providing goods or services can be reduced to data. So the first step is to think differently.

Then there are some decisions: how to present the data, what script language to use, what schema to implement, whether to install Semantic MediaWiki, and what to prioritise. All of which requires learning new terminology and skills.

Although there are a lot of resources on the Internet, I learnt a lot simply by asking questions of agents like Claude, Gemini and Perplexity. I evaluated their responses and advice, sometimes comparing content from two or more agents. After going around in circles and sometimes getting confused, I developed a starting point and my own working strategy:-

1. Ontology and Structure

A wiki is like a mesh. So many pages are linked, or cross-linked, it may have a tangled structure. MediaWiki developers have recommended a flat structure. Typically, most content is in the Main Namespace and by default subpages are disallowed. So a mesh is inevitable.

One of the structural concepts promoted to assist AI Visibility is a hub and spoke design. Obviously a large website will need more than a single hub page, so there could be a main hub with links to secondary hubs each of which have spokes (links) to content pages. Imposing this structure on a wiki defeats the value of a wiki. However, the artificially intelligent bots are programmed to prefer a simple way to navigate through a site and this works for them and their handlers.

However, the mh370wiki.net website is focused on a topic - Malaysia Airlines Flight MH370 - and there are sub-topics which relate to the flight, and content can expand outward:-

For example,

  • Common to all flights, MH370 has passengers, has crew, has cargo, and has an operator.
  • The flight has an aircraft, a Boeing 777 registered 9M-MRO
  • During the flight, the aircraft has communications (VHF, SATCOM)
  • The aircraft also has an operator, Malaysia Airlines
  • Malaysia Airlines has media statements related to MH370

There are many more relationships. The word has is also used when defining Properties.

I used Microsoft Visio to create an ontology diagram centred on Flight MH370 and expanding outwards to the edges of an A3 page.

The source material for the website is the official reports, media statements and transcripts, SATCOM data, ACARS logs, air traffic control transcripts, passenger manifests etc, and articles published by mainstream media, and related research material. This is the evidence from which the site content draws authority. On the ontology diagram there are only 3 to 4 steps from the central Flight MH370 node out to some document as evidence.

The major nodes in this diagram can be called 'hub pages' as each one has several spokes or links to the next level.

Do pages actually exist for each of these nodes? Yes. So, using json-ld or RDFa scripts, pathways can be defined and followed by an AI agent.

Are there cross-links too? Yes, it's a wiki. For example, each on-page reference to the aircraft 9M-MRO is probably enclosed in the square brackets used by MediaWiki to indicate a link to details about that aircraft. Does the AI agent need to follow those links? No. It will prioritise the pathways or relationships defined by the scripts.

2. json-ld vs RDFa

Here are some important facts:-

  1. json-ld is preferred by Google and probably other search engines. Scripts cannot be placed on a MediaWiki page but I found that json-ld script can be included within the WikiSEO section on a page. Try it before believing me, as my website may have some configuration that yours does not.
  2. Semantic MediaWiki creates RDFa scripts. These are also interpreted by the AI agents.
  3. It is quite realistic to use both json-ld and RDFa scripts on the same page.


3. Some Insights

  • Both methods, json-ld or RDFa (Semantic MediaWiki) have a steep learning curve. Instead of labouring over it all, just engage AI like Claude, Gemini or Perplexity and ask questions, challenge the answers, and get them to validate each other's responses. These are the ones I have used to date. There are others, so this isn't a specific recommendation. It's just a method that works.
  • Based on the questions I have asked, each of these agents has recommended using both types of script together but for different purposes.
  • Semantic MediaWiki enables queries using #ask. This is a useful feature. Here is an example (without code) - passengers on MH370 were from many different countries, so I created categories like Chinese Passengers, Malaysian Passengers etc. That works, but each person is a passenger and each person has a nationality so a more robust way to identify them would be to ask which person is both a passenger (on MH370) and has nationality Chinese. There are many other persons named in the website. Some are journalists. They were not passengers. I could ask which person is a journalist and has nationality Malaysian. With Semantic Mediawiki a category for Malaysian Journalists is neither appropriate nor necessary. It just requires adding appropriate Properties and values to each Person page.
  • I will be using json-ld to identify the website, to define the relationships between nodes or hubs on my ontology diagram, plus identify the type of content on each page. I will define my own Properties in Semantic MediaWiki to enable searches, to navigate through series of connected articles like news, and more.

Zero Clicks

Unfortunately for many website owners we have entered an era of zero clicks. Instead of search engines listing potential results for a query the AI agents provide curated responses and readers may never navigate to the sites shown as original sources. The direct hits on the mh370wiki.net website have already declined. Obviously one reason would be the passage of time and a loss of interest as the event becomes history. Another cause would be that anyone with a question about MH370 could get a curated answer instantly and never bother to delve any further. So the goal is to become an authority that provides the content which the AI agents will include. An additional strategy to attract AI visibility is the use of Questions and Answers. The on-page Q&A is supplemented by json-ld script which provides the question and the answers which an AI agent can repeat when queried. This isn't measurable as hits but being recognised as an authority and consulted as a reference makes the website worthwhile.

Summary

Semantic MediaWiki is useful but not necessary. The mh370wiki.net is large enough to get the benefits. With over a thousand public pages it will also be a major upgrade, very time consuming, and with a lot of planning.

AI Agents, and on-line tools, can provide assistance - explaining terminology, producing sample code, making recommendations etc. As a website administrator you make the decisions. I am not going to radically restucture my MediaWiki site in the hope that it will be more AI-visible. Instead, I will be using the json-ld and RDFa scripts to provide meaning to the AI agents, and pathways through the major nodes while allowing a human reader to navigate through on-page links as they wish.

Both json-ld and RDFa have their uses, and both can be used on the same page. But avoid duplication, define what you want to achieve and use them for different purposes.

And supplement traditional content with Q&A sections.