TE — transposable element.
LTR — long terminal repeat. LTR retrotransposons have direct repeats at both ends.
TIR — terminal inverted repeat. DNA TIR transposons have inverted repeats at their ends which are recognized by transposase.
RMSD — root-mean-square deviation, a measure of the average distance between corresponding atoms after two molecular structures are superimposed. Lower values indicate a closer structural match.
Class I transposons, or copy-and-paste TEs — mobile elements that transpose through an RNA intermediate. Their RNA is reverse-transcribed into DNA, which is then integrated as a new genomic copy.
Class II transposons, or DNA transposons — mobile elements that transpose directly as DNA. Cut-and-paste DNA transposons are excised from one genomic site and inserted into another by a transposase.
Autonomous elements encode the protein machinery required for their own transposition, whereas non-autonomous elements depend on machinery encoded by another element.
Last year, I wrote a blog post about the origins of Drosophila melanogaster transposon names and, to my great satisfaction, it found its readers. I promised to write a second part about TEs named after their ability to move (rover, pogo, Nomad, and so on), but, to my own surprise, I ended up starting a project on one of them which took a lot of effort and most of my summer. My group leader and I recently published a preprint with a tongue-in-cheek title that he wasn’t a huge fan of: The Little Transib That Could – and Could Not: Contrasting Invasion Outcomes within a Novel DDE DNA Transposon Lineage. Here, I want to describe the work beyond the paper and highlight what I found most exciting about it. There is a lot to discuss, and I don’t want to turn this into an extremely long read, so I will split it into two or three parts.
What kick-started this project was a simple idea: to look at the distribution of known TEs across Drosophila population data. I already had some experience analysing the TE content of long-read genome assemblies from different insect species, so this did not seem particularly difficult. For this, I picked a particularly nice panel: 14 high-quality PacBio genome assemblies from founder strains of the Drosophila Synthetic Population Resource, a set of consensus TE sequences, Earl Grey from Toby Baril, and a cup or two of Earl Grey tea, which I also like.
I focused on a particular subset of insertions: full-length copies that were unique to a single genome and located outside TE-dense regions, as a rough way of enriching for euchromatic insertions. The results were interesting for four TE families. First, I noticed that roo had a particularly large number of non-shared insertions across the strains. I think this is amazing, and here is my reasoning.
roo is a very interesting LTR retrotransposon. First, it has a complex structure, including a long and structurally variable 5′ UTR with tandem repeats and a microsatellite region. Second, it appears to use early embryogenesis as a transcriptional niche. A spectacular burst of roo expression occurs in the embryonic soma: its transcripts first appear predominantly in yolk nuclei and later in mesodermal cells. At its peak, 4–6 hours after egg laying, roo accounts for more than 1% of the entire embryonic transcriptome and over 70% of all TE-derived reads. I want to highlight that this happens in wild-type flies with piRNA pathway intact. The authors proposed that roo may use this burst in the soma to evade germline silencing, with virion-like particles potentially providing a route back into germ-cell precursors. That last step remains a hypothesis, but the lifestyle itself is remarkable. Together with my observation of numerous strain-specific insertions, and with direct evidence that multiple structural variants of roo remain mobile, this is consistent with ongoing transpositional activity. I think roo is one of the most peculiar eukaryotic transposons ever described, and it definitely deserves further investigation.
Switching to the other TEs, three families showed striking expansions in single strains. The A5 genome had an excess of copia insertions; B1 had plenty of Doc, a LINE-like retrotransposon; and A6 had far too many Hopper copies. Hopper is a roughly 1.4-kb DNA transposon with terminal inverted repeats at its ends. The interesting thing about it is that it is non-autonomous: it may retain the sequences needed for recognition by a transposase but does not encode the transposase itself.
I realised this by examining Hopper insertions with TE+Aid, before I read the Kapitonov and Jurka paper from their “Molecular Paleontology” series that introduced the Transib superfamily to the world. TE+Aid is software that I have run at least several hundred times over the past three years, and I am extremely proud to have made a small contribution to it. Apparently, it is now getting a Claude-assisted rewrite in Python and some updates.

So, what is the first thing we think when we see an expansion of non-autonomous DNA transposons? There should be something moving the thing! And indeed there was. I basically went through the A6 genome in IGV, looking for suspicious matches. I reasoned that the autonomous element should be longer, and luckily I found a sequence with Hopper-matching regions at both ends and a “gap” in the middle. The total length of the guy was 2,847 bp, almost exactly twice that of the short Hopper elements.
I have already mentioned that I used Earl Grey on several occasions. What I have not mentioned is that I had run both Earl Grey and another TE-annotation tool, EDTA, on many different genomes, usually insect genomes, and then gone hunting. The hunt usually involved extracting novel TE sequences, using BLAST to look for evidence of horizontal transfer, and searching for conserved protein domains with the Conserved Domain Database to validate the proposed TE structure.
My long Hopper contained a 648-aa open reading frame that, surprisingly, returned no matches in CDD. This was quite a shock. With endogenous retroviruses, for example, I would not be surprised to find only weak matches to some gag or env proteins, since these are often the most rapidly evolving parts. For the reverse-transcriptase and integrase core, however, I would always expect a match, and that is what I had seen again and again. Here I had a DNA transposon encoding essentially one enzyme, responsible for excising the element and inserting it elsewhere, so I expected to find a conserved transposase domain related to other members of the Transib superfamily.
What made this even wilder was that Transib transposases are distant relatives of RAG1, the catalytic protein at the heart of V(D)J recombination; indeed, the RAG1 core is thought to have evolved from a domesticated Transib-like transposase. I therefore assumed that the superfamily would be decently annotated. Alas.
The name Hopper had been used for the Drosophila element since 1995. However, it also created a naming collision with an unrelated hAT-family element from Bactrocera dorsalis, which was also named hopper in 1997. Rather confusingly, the Bactrocera hopper subsequently acquired a much more extensive element-specific literature. The original defective copy was followed by the identification of an apparently intact candidate in 2003, its successful use as a vector for insect germline transformation in 2020, and an analysis of its distribution across Bactrocera and related tephritid species in 2024. The Drosophila Hopper, by contrast, remained poorly characterised and was known only in its short, non-autonomous form.
The discovery of its autonomous counterpart therefore gave us two reasons to revise the nomenclature: to avoid confusion with Bactrocera hopper and to distinguish two structurally distinct forms of the Drosophila element. Importantly, the short copies were not simply a miscellaneous collection of degraded, non-autonomous insertions. Most shared a characteristic central deletion, defining a coherent structural form that warranted a name of its own. I will return to the origin and significance of this deletion later. Because Nozomi belongs to the Transib superfamily, we decided to continue its railway theme: Transib itself was named after the Trans-Siberian Express. We named the 2,847-bp autonomous element Nozomi and its roughly half-length, non-autonomous derivative Kodama, after two Shinkansen services that operate on the same railway lines. Nozomi is the fastest, limited-stop service, whereas Kodama stops at every station. On the Sanyo Shinkansen, Nozomi services use 16-car trains, while some Kodama services use compact eight-car sets: a nice parallel to our long and short elements.

The analogy is not perfect: both real trains have their own engines, and a Nozomi does not pull a Kodama. In the biological system, Nozomi provides the transposase that mobilises both itself and Kodama. Kodama retains the terminal sequences recognised by this transposase but has lost the machinery needed to move independently. Still, the names capture the part of the analogy that matters: two closely related forms travelling along the same route, one twice as long as the other. And that is how Hopper became two trains.
Having found Nozomi, I used AlphaFold 3 to model its transposase structure. A model containing a transposase dimer, both TIRs and two Mg²⁺ ions positioned the DDE residues in a catalytic centre consistent with the characteristic fold of Transib transposases. I then compared the predicted Nozomi transposase–DNA complex with the experimentally determined strand-transfer complex of HzTransib, a well-characterised Transib element from Helicoverpa zea. Although their N-terminal regions differed substantially, their catalytic cores aligned with an RMSD of only 1.54 Å. When the comparison was restricted to the catalytic residues, the RMSD fell to 0.23 Å, showing that their spatial arrangement was nearly identical. The enzyme that had returned no conserved-domain match was, structurally, unmistakably a Transib transposase.
Finding Nozomi explained what could mobilise Kodama, but it did not explain how Kodama arose or why it became so abundant. That is where I will pick up the story in Part II: I will look more closely at the central deletions found in this lineage, what their breakpoints suggest about how they formed, and the subsequent amplification history of Kodama.
