Biosemiotic knowledge-graph embedding · European ecology

Runa‑2

A geometric map of who relates to whom — the shared substrate beneath nine ecological agents.

WhatA trained BoxE embedding over ~422,000 cleaned + agent-verified ecological relationships
StatusTrained, upgraded & deposited · conventional layer validated · two sign-axes (temperature, pH) independently validated · v2.0 fixes directional scoring (directionality 0.53→0.98)
Built onPyKEEN · open European data
Depositedconcept DOI 10.5281/zenodo.20630499 · v2.0 10.5281/zenodo.21182367
Last updated2026-07-04 (v2.0)
scroll · plain language first
§01 — What it is, plainly

Decades of ecological records, folded into a single space.

Runa-2 takes open European ecological data and turns it into one geometric map of who relates to whom. Every species, and every environmental state, becomes a point in a shared space. Two points sit close together when the things they describe are ecologically related.

A predator lands near its prey’s neighbourhood. A pollinator settles near the plants it visits. A nitrogen-loving plant drifts toward the marker for high-nitrogen soil. Nothing is told where to go — the positions fall out of the relationships themselves.

It learns from typed triples: plain statements of the form (head, relation, tail). Given roughly 490,000 of them, the model arranges every point so that each relation becomes a consistent geometric move. Once that’s done, you can ask questions the source records never stated outright:

one training statement
Bombus terrestris pollinates Salix caprea
~490,000 statements like this become rotations in one shared geometry.
nearest-neighbour
“What is ecologically closest to this species?” — read straight off proximity.
relational
“Who might this species pollinate?” — follow the relation as a geometric operation, head ∘ relation → tail.
gap-filling
Plausible links nobody recorded, suggested by analogy to better-documented systems.

It is not a language model and not a forecaster. It is a knowledge-graph embedding (BoxE, trained with PyKEEN) — closer to a map than to a chatbot.

You shall know a species by the company it keeps.after J. R. Firth, 1957
§02 — Why “biosemiotic”

Species relate not only by contact, but by the signs they share.

On top of the ordinary ecological links — predation, pollination, mycorrhizae — Runa adds sign-relations: species treated as producers and readers of ecological signs within an Umwelt, the perceived world of an organism.

The novel move is small but consequential. These sign-relations are trained as ordinary typed edges, so the model places them in the same geometry as everything else. Two species that read the same environmental sign are pulled together — even if they never directly interact. That “relatedness by shared sign” is what the project is really about.

The intellectual lineage is explicit: Peirce’s index, Uexküll’s function-circle, Hoffmeyer’s semiotic scaffolding. Runa doesn’t try to decode any organism’s inner meaning; it models the relational structure of ecological sign-making, and leaves the inner world untouched.

§03 — At a glance

The artifact, in numbers.

421,563
typed triples
(cleaned + agent-verified)
53,099
entities (+ signal +
navigation)
0.98
directionality
(fwd>rev, v2.0)
0.576
independent thermal ρ
(vs real climate)
0.118
held-out MRR
(leak-free split)
128d
BoxE embedding
dimension
10
countries of
geographic scope

The production model is BoxE on the GBIF-reconciled master. Beyond held-out link prediction, its indicatorOf geometry is now independently validated against real species-occurrence climate (temperature) and soil pH — see §09–§10. v1.4 extends the signal-web layer: which species cast which sounds (song, call, alarm, echolocation), which read the elements (magnetic & electric fields), and how they navigate — by magnetism, the sun, the stars or smell. v2.0 corrects directional scoring: it drops the synthesised inverse triples that had symmetrised the geometry, so a reversed trophic triple (perch “eating” pike) no longer scores as high as the correct one — directionality rose 0.53 → 0.98 with no ranking loss. Published: v2.0 DOI 10.5281/zenodo.21182367.

§04 — Entities & relations

Two kinds of node, two layers of relation.

The nodes

Biological entities — taxa, keyed by scientific name (canonical binomials from GLOBI and Mangal, reconciled to the GBIF backbone — 74,147→65,272 entities after merging synonyms). State & community nodes — the discrete non-biological targets the sign-relations point at: environmental states like soilNitrogen_b4 or thermalIndicator_b3, and detected communities community_0 … community_89.

Layer 1 — conventional ecology

Around 40 relation types, mapped to the OBO Relations Ontology where possible (PURLs taken from GLOBI’s authoritative ro.tsv): trophic (eats, preysOn), antagonistic (parasiteOf, pathogenOf), mutualistic (pollinates, ectomycorrhizalHostOf), structural (createsHabitatFor, epiphyteOf) and generic (interactsWith, coOccursWith).

Layer 2 — biosemiotic (the novel contribution)

ENVAI-namespaced sign-relations
RelationMeaningTail node
indicatorOfspecies is a sign of an environmental state (Peircean index)env-state
keystoneSignProducerInkeystone sign-producer within a community (Hoffmeyer scaffolding)community
perceivesSignalperceptual boundary of the Umwelt (Uexküll) — not builtsignal
phenologicalIndicatorOfphenological sign — defined, not populatedphenophase

v1.4 set PyKEEN’s create_inverse_triples=True, which synthesised an inverse relation for every triple at training time. That symmetrised the per-relation head/tail boxes and destroyed directional scoring — a reversed trophic triple scored as high as the correct one (pike-eats-perch ≈ perch-eats-pike). v2.0 sets create_inverse_triples=False: one canonical direction per relation, no synthesised inverses. Directionality (fraction of asymmetric triples where the correct direction outscores its reverse) rose 0.53 → 0.98 with no ranking loss. Inverse GLOBI labels (eatenBy, pollinatedBy) are still flipped to canonical during ingestion.

§04+ — Structural fragility

What holds each ecosystem together — and where it breaks first.

Composition tells you what is present; network topology tells you what is load-bearing. RUNA reads each ecosystem's interaction graph by k-core shell (how deep a species sits in the structural core) and a cascade simulation (how many species fall with it). The deepest dependency is rarely the apex predator — it is a copepod (Calanus finmarchicus), the Mývatn lake midge, a foundation tree, or a river mussel.

1 graph
Knowledge graph
Species + interactions in Neo4j.
2 model
BoxE embedding
Who-meets-whom across ~46k species.
3 reconstruct
Link prediction
Infers missing wiring for data-poor species.
4 structure
k-core + cascade
Coreness + secondary-extinction impact.
5 read
Load-bearing
What to protect; where it fails first.
Firm: k-core and cascade are deterministic and reproducible from the graph; every load-bearing pick is ecologically defensible. In progress: these run on the sparse curated graph (cores reach 2–3); the richer version densifies each ecosystem with RUNA's predicted links first. Validation: RUNA's indicator placements are checked independently against real GBIF occurrence × climate/soil, not held-out edges.
§05 — Deriving the sign-edges

The signs aren’t in any database. They’re derived — so the derivation is stated, reproducible, and hash-frozen at deposit.

indicatorOf← EIVE / Ellenberg
Each European vascular plant’s published indicator value on six gradients — Light, Temperature, Moisture, Reaction/pH, Nutrients/N, Salinity — binned into five fixed classes over 0–10. These are published indicator values (expert-harmonised rankings), not raw site measurements: legitimate “measured signs” in the bioindication tradition. 3,055 of 8,908 EIVE plants matched existing graph species, wiring the edges into the interaction web.
keystoneSignProducerIn← graph centrality
No usable external keystone dataset exists. Keystone-ness is operationalised as within-community degree centrality above the 98th percentile on the real interaction graph (Louvain communities). Honest framing: “topologically central within a detected community” — a defensible proxy, not an experimentally demonstrated keystone.
perceivesSignal← deliberately deferred
Perception, not production. No species→perceived-signal dataset exists at scale — only slivers (bird UV-vision from opsin genetics; ~34 species in the Animal Audiograms Database). Xeno-Canto documents signal production, which is not a proxy for perception. Rather than fabricate edges, this relation is left for a future, honestly-scoped perception-only build.
open methodological choice The Notion spec describes an alternative indicatorOf route — Living Planet Index population trends × CHELSA climate shifts. EIVE is the niche / literature-attested route; the two are complementary, and the choice is held open (see §10).
§06 — Data & provenance

Open sources, filtered to a northern-European window.

Sources feeding the stage-3 master
SourceRoleLicenseIn-scope
GLOBI 2026-05-29typed interaction triples (the spine)CC0 / CC-BY per dataset368,983
Mangal v2 APIwhole-network structureopen · per-network DOI79,021
EIVE 1.0Ellenberg values → indicatorOfCC-BY 4.044,590
graph centrality (internal)keystoneSignProducerInderived302
GBIFoccurrence + taxonomy backboneper-record CC0 / CC-BYstreamed

EIVE via Zenodo 10.5281/zenodo.7427088. GBIF: 69 GB zip (~300M+ occurrence rows; 10-country, coords, CC0 / CC-BY) is now streamed for the independent thermal-niche validation (§09). Full co-occurrence and environmental-layer use still deferred.

SENOFIDKISNLDEFRBECH

All nine ENVAI agents are covered, France/Ondine included. GLOBI is filtered by a coarse W-Europe/Nordic bounding box (lat 41–71.5, lon −25–32), to be refined by reverse-geocoding.

§07 — Pipeline (ETL)

From raw records to one master triple file.

Each stage is a module under src/envai/; every triple file carries [head, relation, tail, source, confidence].

  1. stage1_globiStream GLOBI, filter to scope, map raw interaction names to the canonical vocabulary (collapsing inverses), dedupe.
  2. stage3_mangalWalk the Mangal API (network / node / interaction), map types to the vocabulary.
  3. stage_indicatorParse EIVE, bin each 0–10 gradient into 5 classes, match plants to graph entities, emit (plant, indicatorOf, gradient_bin).
  4. stage_keystoneBuild the interaction graph, run Louvain community detection, take the top 2% within-community degree → (species, keystoneSignProducerIn, community).
  5. stage5_taxonomyConcatenate and dedupe sources into a master, then reconcile names onto the GBIF backbone — 14,172 names merged (74,147→65,272 entities).
  6. train · eval_per_relationRotatE / ComplEx via PyKEEN, then a per-relation metric breakdown.
  7. serve_neo4jPush embeddings to a Neo4j vector index for queries — built, not yet run.

Deferred but coded: stage4_env (CHELSA / LUCAS → respondsTo* nodes), and serving. stage2_gbif extraction had a read-all-into-RAM bug that OOM-killed on the 100 GB CSV — fixed to stream with shutil.copyfileobj.

Independent validation path: scripts/extract_niche_multi.py streams the GBIF zip into per-species realized niches across several environmental rasters (CHELSA temperature & precipitation, SoilGrids pH & nitrogen), then envai.validate_indicator Spearman-correlates the model's indicator placement against the real niche, per gradient. The EIVE-value-vs-reality column doubles as a proxy-validity check.

§08 — Embedding model & training

BoxE, tuned to run on one Blackwell card.

Training configuration
ModelBoxE, embedding_dim 128 — upgraded from the initial RotatE (see §09); box embeddings handle the 1-to-N sign-relations RotatE could not.
LossNSSALoss (self-adversarial), margin 9.0, adv. temperature 1.0
OptimizerAdam, lr 1e-3
LoopsLCWA, basic negative sampler, 8 negatives/positive, batch 2048
Epochsv2.0: 150, on all triples (production model). v1.4 used 300 with early stopping (freq 20, patience 5, metric = MRR).
Splitv2.0: inverse triples OFF; production model trained on all 337,561 triples. Evaluation on a leak-free split by unordered entity-pair — a random 80/10/10 split leaks reversed pairs, which both inflates MRR and hides direction-blindness. (v1.4: 80/10/10 random, inverse triples on.)
EvalRankBasedEvaluator, bounded (batch 128, slice 2048) to dodge the WDDM watchdog
ComparatorsComplEx underperformed RotatE everywhere (MRR 0.177 vs 0.283). BoxE (a box embedding) was then adopted as the production model — it lifted the independent thermal validation from ρ 0.114 to 0.545 while improving link prediction (see §09).

Training the ~490K-triple / 74K-entity graph at dim 128 takes roughly 15–35 minutes on a single modern GPU (~2–5 s/epoch).

§09 — Results

What the geometry learned — and where it’s honest about plateauing.

Model snapshots · filtered ranking. Rows through v1.4 use a random 80/10/10 split (MRR inflated by reverse-pair leakage); v2.0 is measured on a leak-free split by entity-pair.
SnapshotTriplesEntitiesRel.MRRH@1H@3H@10Dir.
stage1 — GLOBI only368,98358,167210.26450.1670.2910.465
stage2 — +Mangal444,97968,183250.28240.1820.3140.487
stage3 — +biosemiotic (RotatE)489,87174,147270.28270.1840.3150.484
BoxE — same data, box model489,87174,147270.28770.1900.3210.487
BoxE + reconciliation — v1.1475,31165,272270.28840.1920.3210.482
BoxE + cleaned + agent-verified — v1.2421,56353,088300.30470.2050.3440.505
BoxE + signal + navigation — v1.4 (inverse=on, random split)421,95253,099330.30100.2010.3390.5010.53
BoxE inverse=False — production v2.0 (leak-free)337,56153,099330.1180.0500.1090.2620.98

The two right-most numbers tell the real story. v1.4's 0.30 MRR was a random-split artifact: reversing pairs leaked between train and test, inflating the score and masking the flaw. On an honest leak-free split (split by unordered entity-pair) v1.4 scores 0.114 MRR and, crucially, only 0.53 directionality — barely above chance, meaning it scored reversed trophic triples as high as correct ones. v2.0 (inverse triples off) lifts directionality to 0.98 at the same honest MRR (0.118); RotatE and PairRE were tried and did not beat it, so the ~0.12 MRR is a data ceiling, not an architecture limit. (Triple counts: the typed graph holds ~422K cleaned triples; the model trains on 337,561 numeric training triples — the same edges after collapsing to canonical direction.)

Per-relation MRR (stage-3, RotatE)

biosemiotic relation conventional relation

indicatorOf is placed in the geometry more consistently than pollination, symbiosis, mutualism and mycorrhizae — it is genuinely learned. keystone scores highest, but on only 17 test edges.

read this before trusting the bars Both biosemiotic relations have small tail vocabularies (30 env-states, 90 communities), so tail-prediction is easy and inflates their scores relative to the 74K-entity species direction. The numbers show the relations are learnable — not that they capture independent ecological truth.

Independent validation — non-circular (the headline)

For each EIVE gradient with a real-world proxy, the model's indicatorOf placement is Spearman-correlated with each species' realized niche from real GBIF occurrences × real environment — independent of the EIVE values it trained on. The expert-vs-reality column is both the ceiling and a proxy-validity check.

Model indicator placement vs real occurrence environment · BoxE + reconciliation · ~5,000 plants
EIVE gradientreal proxymodel ρexpert ceiling ρverdict
TemperatureCHELSA bio10.5760.797validated
Reaction (pH)SoilGrids pH0.3160.213validated — above ceiling
NutrientsSoilGrids total N−0.10−0.20proxy invalid
MoistureCHELSA precip0.070.09proxy invalid

Two axes — temperature and pH — are independently validated. On pH the embedding even beats the expert ceiling (0.335 vs 0.213): it denoises EIVE's per-species values across the interaction graph. Moisture and nutrients fail at the proxy level — the expert-vs-reality ceiling is itself ≈0, so annual precipitation and total soil N aren't valid stand-ins for site-moisture / nutrient-availability. A built-in self-check, not a model failure.

§10 — Validation status & honest limitations

This section is deliberately strict. It’s what separates an honest artifact from an overclaim.

Independent validation now exists for two axes (§09). Temperature (ρ 0.545) and pH (ρ 0.335) are validated non-circularly against real occurrence × environment — independent of the EIVE rule the edges were derived from; on pH the model even beats the expert ceiling. Still unvalidated: moisture & nutrients (no valid proxy found yet) and keystoneSignProducerIn, which still rests on a within-rule check.
The 1-to-N weakness was fixed by switching to BoxE. RotatE under-encoded the 1-to-N sign-relations (thermal ρ 0.114); BoxE — a box embedding built for 1-to-N — lifted that to 0.545. Taxonomy stays a structural constraint, never scored triples.
Small tail vocabularies inflate the biosemiotic metrics — see the §09 caveat.
GBIF taxonomy reconciliation is done. Names are merged onto the GBIF backbone — 14,172 collapsed, 74,147→65,272 entities — so synonyms (e.g. Phragmites communisP. australis) now share a node.
Coverage gaps. GBIF co-occurrence and the CHELSA / LUCAS environmental layer are coded but not built; the bounding-box filter is coarse; counts are point-in-time.
Sign-edges are interpretive by construction — hypotheses the embedding helps test, documented with provenance, not ground truth.
v1.2 — Built in the open

The v1.2 gains came from nine place-based agents — and every edge was triple-checked.

Nine ecological agnt eco agents each proposed dependency edges from the place they inhabit, each with a primary-literature source. Every claim was then fact-checked by a multi-model consensus — the original extraction, then claude-opus-4-8 and gpt-5.4 reading the cited sources. An edge entered the model only where Claude confirmed it and GPT did not refute it. 127 proposed → 96 entered → 39 at full ≥0.85 three-model consensus. The 9 that GPT refuted — including a reversed oak/beetle edge it caught at ρ 0.18 — were dropped.

A sample of agent-verified edges · every fact public, hashed and voteable
agentedgefact_hashClaudeGPT-5.4consensus
aegirLepeophtheirus salmonis parasitises Salmo salar72b74120c5f0b25c0.96
eldvatnSalmo trutta feeds on Simulium vittatum547232729f9190460.95
alvaAphanomyces astaci parasitises Astacus astacus9a6298b217c4e98e0.95
aegirGadus morhua feeds on Mallotus villosusc5329671002e844c0.93
ondineLarus genei feeds on Artemiae34bcb84a9112a150.91
scîrwuduThaumetopoea processionea feeds on Quercus robur7fdac38602d2b8430.91
alvaUnio crassus depends on Phoxinus phoxinus83e7b4dd1cd6cfa90.90
ondineTuber magnatum mycorrhizal with Quercus pubescens7dabba532d17c06f0.89

Nothing here is a black box: the full evidence reports are public on the community forum, every vote and its notes live in the fact-verification system, and the frozen, hashed edge set ships in the Zenodo deposit.

§11 — The priority claim

The first trained relational embedding that operationalizes biosemiotic relations for ecology.

Narrow and defensible — not “first biosemiotic AI” in general. Here is exactly what is true today:

The embedding exists, is trained, and is on disk — this resolves the “model vs. plan” question.
The conventional-ecology embedding is validated by held-out link prediction.
Two biosemiotic relations are in the trained geometry and learnable.
Two sign-axes (temperature, pH) are now independently validated against real occurrence × environment (§09); moisture / nutrients await valid proxies.
perceivesSignal, one of the three named relations, is not built.
for an honest deposit The novelty lives in the relations, not the weights. Deposited on Zenodo — DOI 10.5281/zenodo.20630499 — with a hash-frozen snapshot of the exact triple set, the ETL / derivation scripts, the model and the schema. Honest label: a validated ecological KGE whose biosemiotic layer is independently validated on two axes (temperature, pH).
§12 — Roadmap

In priority order.

  1. Independent (non-circular) eval — done for 2 axes ✓Temperature (ρ 0.576) and pH (ρ 0.316) validated against real occurrence × environment. Remaining: valid proxies for moisture & nutrients, and an LPI × CHELSA population-trend axis.
  2. GBIF taxonomy reconciliation — done ✓14,172 names merged onto the GBIF backbone (74,147→65,272 entities).
  3. Box / hyperbolic comparator — done ✓BoxE adopted as the production model; it fixed the 1-to-N weakness (thermal ρ 0.114→0.545).
  4. GBIF co-occurrence + CHELSA / LUCASThe respondsTo* environmental layer (occurrence already downloaded).
  5. ServingNeo4j vector index with nearest-neighbour, analogical and hybrid queries for the nine agents.
  6. HPO & a perceivesSignal buildPerception-only, from audiogram / opsin data.
  7. Deposit — done ✓Hash-frozen triples + scripts + schema + model on Zenodo: DOI 10.5281/zenodo.20630499.