Skip to main navigation Skip to search Skip to main content

A systematic comparison of transformers and ConvNets for root segmentation across nine datasets

  • Abraham George Smith
  • , Sotiris Lamprinidis
  • , Anand Seethepalli
  • , Larry M. York
  • , Eusun Han
  • , Patrick Möhl
  • , Kyriaki Boulata
  • , Kristian Thorup-Kristensen
  • , Jens Petersen

Research output: Contribution to journalArticlepeer-review

Abstract

Background: Root segmentation is a fundamental yet challenging task in image-based plant phenotyping. Accurate segmentation is a prerequisite for extracting root traits relevant to plant physiology, breeding, and agronomy. While U-Net and other convolutional neural network (ConvNet) architectures have been applied to root segmentation, no systematic comparison of multiple Transformer and ConvNet architectures has been conducted across diverse root imaging conditions. Results: We evaluated 21 segmentation architectures across nine diverse root image datasets, training 1511 models to assess all combinations of architecture, dataset, pre-training strategy, and learning rate, producing over 3 million segmentations for evaluation. Transformer-based models significantly outperformed ConvNets for Dice (mean Dice 0.679 vs 0.659;). Root-diameter and root-length correlation were also higher for Transformers, but the differences were not statistically significant (and respectively). Pre-training significantly improved mean Dice from 0.623 to 0.666 (), with Transformers benefiting more from pre-training than ConvNets (Dice improvement + 0.072 vs + 0.021;), supporting the hypothesis that fine-tuned Transformers transfer more effectively across large domain gaps. MobileSAM achieved the highest Dice score (0.693) while maintaining computational efficiency. Both architecture families underestimated thin root length compared to manual annotations. Dataset choice explained 70.9% of performance variance, far exceeding model architecture (6.7%). Purpose: Transformer architectures significantly outperform ConvNets for root segmentation accuracy, and pre-training significantly improves performance, particularly for Transformers. Pre-trained MobileSAM offers the best accuracy at competitive computational cost. Dataset choice dominates performance variance, suggesting practitioners should prioritize data curation over architecture selection.

Original languageEnglish
Article number52
JournalPlant Methods
Volume22
Issue number1
DOIs
StatePublished - Dec 2026

Funding

Open access funding provided by Copenhagen University. This work was supported by Novo Nordisk Foundation grant NNF22OC0080177 for Abraham George Smith and Jens Petersen. This material is based in part upon work at the Center for Bioenergy Innovation supported by the U.S. Department of Energy, Office of Science, Biological and Environmental Research under Contract Number ERKP886. This Thismanuscript has been authored in part by UT-Battelle, LLC, under contract DE-AC05-00OR22725 with the US Department of Energy (DOE). The publisher acknowledges the US government license to provide public access under the DOE Public Access Plan license to provide public access under the DOE Public Access Plan(https://energy.gov/doe-public-access-plan).

Keywords

  • Benchmark
  • Convolutional neural network
  • Deep learning
  • Image analysis
  • Minirhizotron
  • Plant phenotyping
  • Root segmentation
  • Transfer learning
  • Transformer

Fingerprint

Dive into the research topics of 'A systematic comparison of transformers and ConvNets for root segmentation across nine datasets'. Together they form a unique fingerprint.

Cite this