You will be provided with a reference and some statements. Please determine whether each statement is 'supported', 'unsupported', or 'unknown' with respect to the reference. Please note:
First, assess whether the reference contains any valid content. If the reference contains no valid information, such as a 'page not found' message, then all statements should be considered 'unknown'.
If the reference is valid, for a given statement: if the facts or data it contains can be found entirely or partially within the reference, it is considered 'supported' (data accepts rounding); if all facts and data in the statement cannot be found in the reference, it is considered 'unsupported'.

You should return the result in a JSON list format, where each item in the list contains the statement's index and the judgment result, for example:
[
    {
        "idx": 1,
        "result": "supported"
    },
    {
        "idx": 2,
        "result": "unsupported"
    }
]

Below are the reference and statements:
<reference>
The Turning Point of 3D Plant Phenotyping: 3D Foundation Models Enable Minute-to-Second Cross-Crop Reconstruction and Beyond



Report GitHub Issue

×

Title:

Content selection saved. Describe the issue below:

Description:

Submit without GitHub

Submit in GitHub

arXiv is now an independent nonprofit!

Learn more

×

Back to arXiv

Why HTML?

Report Issue

Back to Abstract

Download PDF

Abstract

1
Introduction

2
Materials and methods

2.1
Data acquisition

2.2
3DFM-based rapid initialization

2.3
Geometry-constrained 3DGS densification

2.4
Few-view reconstruction via iterative view synthesis and enhancement

2.5
2D-to-3D semantic transfer

2.6
Scale recovery and leaf instance separation

2.7
Phenotypic measurement

3
Results

3.1
Overall framework

3.2
3DFM-based initialization reduces the plant reconstruction front end from minutes to seconds

3.3
Comparison of performance in few-view reconstruction

3.4
2D-to-3D semantic transfer achieves more consistent plant organ segmentation than direct 3D approaches

3.5
Leaf area and inclination angle are reliably estimated from reconstructed 3D plant structures

3.6
All three design choices improve performance, with sparse initialization contributing the strongest gain

4
Discussion

4.1
From 3DFM-based initialization to measurable phenotyping

4.2
Current limitations and applicability boundaries

4.3
Implications for low-cost and cross-crop plant phenotyping

5
Conclusions

References

License: arXiv.org perpetual non-exclusive license

arXiv:2607.01753v1 [cs.CV] 02 Jul 2026

The Turning Point of 3D Plant Phenotyping: 3D Foundation Models Enable Minute-to-Second Cross-Crop Reconstruction and Beyond

Journal:
Artificial Intelligence in Agriculture

Hanyue Jia

Note:
These authors contributed equally to this work.

Affiliation:
Northwest A&F University, Yangling, 712100, China

Wei Zhou

Note:
These authors contributed equally to this work.

Affiliation:
Huazhong University of Science and Technology, Wuhan, 430074, China

Wenbo Zhou

Affiliation:
Wuhan Institute of Technology, Wuhan, 430205, China

Yanan Li

Affiliation:
Wuhan Institute of Technology, Wuhan, 430205, China

Hao Lu

Email:
hlu@hust.edu.cn

Corresponding author:
Corresponding author.

Affiliation:
Huazhong University of Science and Technology, Wuhan, 430074, China

Tingting Wu

Email:
tt_wu@nwsuaf.edu.cn

Corresponding author:
Corresponding author.

Affiliation:
Northwest A&F University, Yangling, 712100, China

Abstract

3D plant phenotyping is notoriously known to be procedure-complicated and of low throughput due to the extensive multi-view imaging, the fragile 3D reconstruction pipeline, and the additional cost from reconstructed geometry to phenotypic extraction. These limitations are further amplified in low-cost data acquisition, where smartphone videos or sparsely sampled multi-view images provide limited view overlap and self-occlusion. In this work, we show that the conventional 3D plant phenotyping pipeline could be streamlined and significantly accelerated with 3D Foundation Models (3DFMs), and particularly, present one of the first cross-crop 3D phenotyping frameworks powered by 3DFMs. The framework replaces COLMAP-style sparse initialization with 3DFM-based feed-forward geometric recovery, combines geometry-constrained 3D Gaussian Splatting for dense reconstruction, enables few-view reconstruction through iterative view synthesis and refinement, and converts reconstructed geometry into measurable organs through 2D-to-3D semantic transfer, metric scale recovery, and organ instance separation. We further construct a cross-crop dataset with smartphone-based image acquisition, diverse plant morphologies, and manual annotations for segmentation and phenotypic evaluation. Experiments across
26
26
plant sequences show that 3D Foundation Models reduce the average reconstruction time from
6.52
6.52
minutes to
1.58
1.58
seconds while maintaining high reconstruction quality and phenotyping accuracy. These results suggest a fresh technical route for high-throughput 3D plant phenotyping, from low-cost image acquisition to fast reconstruction, perception, scale recovery, and phenotypic measurement.

Keywords:

3D plant phenotyping , 3D Foundation Models , 3D Gaussian Splatting , few-view reconstruction

1
Introduction

Plant phenotype
links between genotype and environment
(
Houle et al., 2010
)
.
To acquire reliable plant phenotypes, extensive sensing technologies have been used, particularly image-based sensing
(
Hu et al., 2025
;
Tao et al., 2022
)
.
Compared with 2D imaging, 3D phenotyping captures canopy architecture, organ geometry, and plant structure more directly,
making it possible to quantify geometric traits such as leaf orientation, light interception, and architectural organization
(
Okura, 2022
;
Yang and others, 2026
)
. As high-throughput phenotyping continues to expand, a central challenge is how to acquire quantitatively usable 3D plant structure at low cost and high efficiency
(
Furbank and Tester, 2011
;
Yang et al., 2020
)
.

Early 3D plant phenotyping relied mainly on active sensing systems, including LiDAR, structured light, and time-of-flight cameras
(
Okura, 2022
)
. These platforms can provide accurate depth measurements and dense point clouds, but their cost, deployment complexity, and downstream processing requirements have limited broad adoption. To lower the hardware barrier, later studies increasingly turned to passive multi-view reconstruction using consumer cameras or smartphones
(
Li et al., 2025
;
Zheng et al., 2026
)
. Structure from Motion (SfM) and Multi-View Stereo (MVS) offered a practical route toward low-cost 3D modeling
(
Schonberger and Frahm, 2016
;
Furukawa and Ponce, 2010
)
, yet these pipelines rely on local feature extraction, cross-view matching, incremental pose estimation, and dense correspondence search, making them computationally time-consuming and strongly dependent on sufficient image overlap. In plant scenes, such dependence is further challenged by repeated texture, severe self-occlusion, and thin structures
(
Harandi et al., 2023
)
. Under sparse-view or rapid-acquisition conditions, classic pipelines often fail during camera initialization or produce unstable geometry, which limits their usefulness for high-throughput phenotyping
(
Schonberger and Frahm, 2016
)
.

Neural rendering has
recently advanced image-based 3D reconstruction
(
Mildenhall et al., 2020
)
. Neural Radiance Fields demonstrated the potential of continuous scene representation, and 3D Gaussian Splatting (3DGS) greatly improved rendering efficiency, becoming a major explicit representation framework
(
Kerbl et al., 2023
)
. Subsequent variants, including SuGaR, 2DGS, and Mip-Splatting, improved surface alignment, geometric consistency, and anti-aliasing
(
Yu et al., 2024
)
.

Plant-oriented studies have also begun to connect neural or point-based 3D representations with phenotypic analysis. PlantSegNeRF
(
Yang and others, 2025
)
explores NeRF/3DGS-based instance extraction, IPENS
(
Song et al., 2025
)
focuses on interactive phenotypic analysis, Wheat3DGS
(
Zhang et al., 2025a
)
investigates crop-specific organ measurement, PlantGaussian
(
Shen et al., 2025
)
extends 3D Gaussian splatting to cross-time and cross-scene plant representation, Hyperspectral Imaging Meets 3D Gaussian Splatting
(
Deng et al., 2026
)
further pushes plant-oriented 3DGS beyond pure morphology toward multimodal structural–spectral representation, and Eff-3DPSeg
(
Luo et al., 2023
)
addresses weakly supervised point-cloud segmentation. These studies show that 3D reconstruction is moving toward plant perception and measurement, but most remain focused on specific crops, organs, or segmentation tasks. More importantly, standard 3DGS still depends on COLMAP-style sparse initialization, so the fragile front end is largely unchanged
(
Kerbl et al., 2023
)
. Its optimization target also remains primarily photometric, whereas plant phenotyping requires geometrically reliable representations of thin leaves, fine stems, and open boundaries for downstream measurement
(
Ojo et al., 2024
)
.

Recently
3D foundation models (3DFMs) emerge and provide a new entry point for overcoming this front-end

bottleneck
El Banani et al. (2024)
.
Their progress is not simply a matter of applying larger models to reconstruction, but reflects a shift in how multi-view geometry is inferred. DUSt3R
(
Wang et al., 2024
)
moved multi-view recovery from feature matching and triangulation toward feed-forward point-map prediction, and MASt3R
(
Leroy et al., 2024
)
further coupled matchability with geometric consistency. VGGT
(
Wang et al., 2025a
)
unified camera parameters, point maps, depth maps, and point tracks within a single model, making it possible to obtain a more complete geometric initialization in one forward pass.
π
3
\pi^{3}

(
Wang et al., 2025b
)
further removed the fixed-reference-view assumption and improved robustness to input ordering through a reference-free, permutation-equivariant design. Recent methods such as FLARE
(
Zhang et al., 2025b
)
and AMB3R
(
Wang and others, 2025
)
extend this trend toward joint recovery of cameras, geometry, and appearance under sparse-view settings.
In short, these methods shift front-end
geometric
recovery from local matching-based incremental optimization to rapid geometric inference driven by learned multi-view priors, creating a practical opportunity for second-level camera initialization and initial structure recovery in plant scenes
(
Xu et al., 2022
)
.

Nevertheless, existing progress has not yet overcome the front-end constraint in low-cost image-based 3D plant phenotyping. Both standard 3DGS and most image-based reconstruction pipelines still depend on COLMAP-style initialization. In plant scenes, its local feature matching and incremental pose estimation are not only time-consuming but frequently become the primary failure point under sparse-view or rapid-acquisition conditions.
In contrast, 3DFMs provide a new opportunity to move beyond this matching-based front end, but their ability to stably replace COLMAP on cross-crop plant data, reduce the number of views required for usable reconstruction, and still support downstream organ-level phenotypic measurement remains insufficiently validated in the plant domain
(
Wang et al., 2025a
;
Wang et al., 2025b
)
. Meanwhile, existing plant-oriented 3D methods have begun to integrate reconstruction with segmentation and phenotypic analysis, yet they often rely on already available 3D results or focus on specific crops, organs, and processing stages
(
Yang and others, 2025
;
Song et al., 2025
;
Zhang et al., 2025a
;
Luo et al., 2023
)
. The key gap therefore is to
explore whether
3DFMs can provide
measurement-ready geometric initialization under low-cost, few-view acquisition conditions and support a measurement-oriented workflow from dense representation, semantic interpretation, and scale recovery to organ-level trait
extraction.

Figure 1
:
Comparison between classical and 3DFM-based pipelines for 3D plant phenotyping.

Classical pipelines typically proceed from image acquisition to COLMAP/SfM-based initialization, dense reconstruction, and manual or interactive trait extraction, and are limited by minute-level initialization, dense-view dependence, and manual intervention.
The proposed pipeline replaces the matching-based front end with 3DFM-based second-level initialization, followed by 3DGS densification, 2D-to-3D segmentation, and organ-level phenotypic extraction.
This design converts low-cost image inputs into measurable 3D plant traits through a faster and more measurement-oriented reconstruction workflow.

In this work, we propose to replace COLMAP-style sparse initialization with
3DFMs for 3D plant phenotyping, achieving second-level recovery of camera parameters and initial structure
(
Wang et al., 2025a
;
Wang et al., 2025b
)
. This initialization is then converted into a dense 3D representation that is better suited for delineating leaf boundaries, fine stems, and canopy structures through geometry-constrained 3DGS
(
Yu et al., 2024
)
. Under few-view conditions, iterative view synthesis and refinement are employed to supplement effective observations
(
Long et al., 2022
;
Liu et al., 2023
)
. Finally, 2D-to-3D semantic transfer, metric scale recovery, and
organ instance separation convert the reconstructed geometry into measurable
organ-level instances. In this way, 3DFMs, 3DGS, 3D perception, and phenotypic computation serve a unified goal:

the rapid generation of measurable 3D plant structures
.

Experiments across different crop species and varying acquisition conditions show that
3DFMs advance the front end of 3D plant phenotyping from the minute level to the second level while preserving the geometric consistency required for organ-level measurement. Across
26
26
plants, the 3DFM-based front-end initialization reduces the average time from
6.52
6.52
minutes to
1.58
1.58
seconds, with comparable reconstruction quality.
Across representative multi-crop samples, the ratio between estimated and measured total leaf area remains within
0.9514
0.9514
–
1.0629
1.0629
, and the mean absolute error of leaf inclination angle is approximately
2.04
∘
2.04^{\circ}
. These results indicate that
3DFMs not only accelerate plant 3D reconstruction, but can also serve as a geometric entry point for low-cost cross-crop phenotyping, supporting a continuous transition from reconstruction to perception and organ-level trait extraction
(
Okura, 2022
)
.

The main contributions of this study
include the following.

1.

We systematically validate 3DFMs as second-level front-end initializers for plant 3D phenotyping.

VGGT and
π
3
\pi^{3}
compress the COLMAP-style front end, including feature matching, pose estimation, and sparse reconstruction, into feed-forward geometric prediction, simplifying and accelerating initialization from minutes to seconds.

2.

We propose a 3DFM–3DGS framework that lowers the view threshold for plant reconstruction.

The framework combines 3DFM initialization, 3DGS densification, and view supplementation, enabling few-view inputs to support usable reconstruction and phenotypic measurement.

3.

We
contributed a cross-crop 3D phenotyping dataset with organ-level ground truths.

The dataset covers diverse crops, morphologies, environments, and growth stages, and provides leaf-instance annotations and manual trait measurements for joint evaluation of reconstruction, perception, and measurement.

2
Materials and methods

2.1
Data acquisition

To support the evaluation of reconstruction efficiency, few-view robustness, 3D semantic segmentation, and organ-level phenotypic measurement, we constructed a smartphone-based cross-crop dataset combining multi-view video acquisition, fixed-budget frame sampling, and manual ground-truth annotation. All sequences were captured using a consumer smartphone with a resolution of 1080
×
\times
1920 and a frame rate of 60.03 FPS under variable illumination. Each video was recorded around a single plant along a continuous closed-loop trajectory to ensure sufficient viewpoint coverage for camera pose estimation and 3D reconstruction. To balance reconstruction completeness against computational cost, a height-adaptive sampling strategy was used, with 100 frames retained for plants taller than 100 cm and 80 frames for plants of 100 cm or less. The overall acquisition protocol and dataset composition are summarized in Fig.
2
.

The dataset comprised a robustness subset and a quantitative evaluation subset. The robustness subset contained 26 sequences named by species, acquisition condition, and sequence index, such as Tobacco_Outdoor_01, Maize_Field_01, and Wheat_Indoor_01. This subset was used to assess performance across species, growth stages, acquisition environments, and structural complexity. It covered maize, tobacco, wheat, soybean, bamboo, rapeseed, legume, pea, broccoli, and related crops under indoor, outdoor potted, and field conditions. Among them, eight sequences were derived from the public dataset Splanting: 3D plant capture with Gaussian splatting
(
Ojo et al., 2024
)
. The quantitative evaluation subset consisted of a mature-plant subset and a temporal maize-seedling subset. The mature-plant subset included five plants from soybean, sesame, and maize, with manual leaf-instance annotations and leaf-level phenotypic ground truth for segmentation and trait evaluation. The temporal subset included nine maize seedlings recorded at four time points, yielding 36 video sequences for stage-wise phenotypic evaluation.

Acquisition trajectories were adjusted according to plant height. Plants taller than 100 cm were recorded using an upward spiral trajectory from the basal stem to the canopy, whereas plants of 100 cm or less were captured using a circular trajectory with an approximately 45
∘
downward viewing angle. In all cases, trajectory continuity and loop closure were maintained to ensure complete plant coverage. For sequences used for absolute phenotypic measurement, the pot diameter was measured before imaging and the pot rim was kept visible to provide an in-scene metric reference for scale recovery.

Manual measurements were collected for plant height, leaf length, leaf width, leaf area, and leaf inclination angle. Leaf length, width, and area were measured using the OpenPheno
(
Hu et al., 2025
)
mini-program, whereas leaf inclination angle was measured using a leaf inclination meter.
To ensure consistency with algorithmic outputs, leaf length was defined along the major axis, leaf width as the maximum transverse width, and leaf area as the one-sided area of a single leaf blade. Leaf inclination angle was defined as the acute angle between the fitted mean leaf plane and the horizontal plane. In the subsequent 3D computation, this quantity is equivalently expressed as the acute angle between the mean leaf surface normal and the global vertical direction.
Repeated measurements were averaged, and the corresponding variance was recorded as an estimate of annotation noise
(
Itakura and Hosoi, 2019
)
.
All experiments were conducted on a Linux workstation equipped with an NVIDIA RTX A6000 GPU with 48 GB memory. The software environments followed the official configurations of VGGT and
π
3
\pi^{3}
, based on Python 3.10 or later and CUDA-enabled PyTorch. VGGT,
π
3
\pi^{3}
, SAM, and Difix3D+ were used with public pretrained weights without further fine-tuning. The 3DGS/Mip-Splatting optimization was performed for 30,000 iterations using default settings unless otherwise specified. Novel-view augmentation was used only in the few-view experiments and ablation studies.

Figure 2
:
Dataset overview and acquisition protocol.
Smartphone videos were acquired along closed-loop trajectories with height-adaptive shooting and frame sampling. The dataset included a robustness subset spanning multiple species and conditions, and a quantitative subset with manual annotations for segmentation and phenotypic evaluation.

2.2
3DFM-based rapid initialization

Front-end initialization recovers camera parameters and initial geometry as the common starting point for densification, semantic transfer, and phenotypic measurement. In plant scenes, it must also provide cross-view consistency in a unified coordinate system. Conventional SfM/COLMAP obtains this through local feature extraction, matching, pose estimation, and sparse reconstruction
(
Schonberger and Frahm, 2016
)
, but repetitive textures, thin organs, occlusions, and uneven views reduce matching stability and increase cost
(
Okura, 2022
;
Harandi et al., 2023
)
. This stage therefore targets rapid recovery of stable cameras and initial geometry.

We use visual-geometry 3D foundation models (3DFMs) as feed-forward initializers. Unlike SfM, which reconstructs cameras and structure from local correspondences, 3DFMs infer multi-view geometry directly from images with learned priors. Following DUSt3R and MASt3R
(
Wang et al., 2024
;
Leroy et al., 2024
)
, we evaluate VGGT and
π
3
\pi^{3}
as representative initializers
(
Wang et al., 2025a
;
Wang et al., 2025b
)
. VGGT predicts camera and geometric attributes jointly, whereas
π
3
\pi^{3}
uses a reference-free, permutation-equivariant formulation. Both share the same conversion interface, and
π
3
\pi^{3}
is used as the default initializer unless otherwise specified.

Let the sampled multi-view images be denoted as

I
=
{
I
i
}
i
=
1
N
,
O
raw
m
=
F
3
​
D
​
F
​
M
m
​
(
I
)
,
m
∈
{
VGGT
,
π
3
}
.
I=\{I_{i}\}_{i=1}^{N},\qquad O_{\mathrm{raw}}^{m}=F_{\mathrm{3DFM}}^{m}(I),\quad m\in\{\mathrm{VGGT},\pi^{3}\}.

Here,
O
raw
m
O_{\mathrm{raw}}^{m}
denotes the model-specific raw geometric outputs, including camera predictions, view-wise geometric maps, depth-related quantities, and confidence or validity information. Because these outputs differ across models,
O
raw
m
O_{\mathrm{raw}}^{m}
serves as a unified notation rather than assuming identical formats for VGGT and
π
3
\pi^{3}
.

The downstream pipeline requires standardized cameras and an initial sparse point cloud rather than raw point or depth maps. We therefore introduce a bridge conversion step:

(
Θ
,
P
s
)
=
F
convert
​
(
O
raw
m
)
,
Θ
=
{
(
K
i
,
T
i
)
}
i
=
1
N
.
(\Theta,P_{s})=F_{\mathrm{convert}}(O_{\mathrm{raw}}^{m}),\qquad\Theta=\{(K_{i},T_{i})\}_{i=1}^{N}.

This conversion reorganizes camera predictions into a unified intrinsic–extrinsic representation, maps model-specific geometry into a common coordinate frame, filters unreliable predictions, and fuses the remaining geometry into
P
s
P_{s}
. The resulting
(
Θ
,
P
s
)
(\Theta,P_{s})
is comparable to COLMAP initialization for 3DGS, but is obtained by feed-forward 3DFM inference rather than matching-based incremental reconstruction. It provides reconstruction-scale initialization, while absolute metric scale is recovered later using the in-scene pot reference.

Through this interface, VGGT and
π
3
\pi^{3}
serve as interchangeable 3DFM front ends. The standardized
(
Θ
,
P
s
)
(\Theta,P_{s})
provides direct geometric input for Mip-Splatting-based densification, 2D-to-3D semantic transfer, and organ-level phenotypic measurement.

2.3
Geometry-constrained 3DGS densification

The 3DFM initialization result
(
Θ
,
P
s
)
(\Theta,P_{s})
provides stable camera parameters and initial sparse geometry, but it is still insufficient for the semantic and geometric operations required by organ-level measurement. For plant scenes, the 3D representation must cover organ structures more completely and preserve geometric continuity around leaf boundaries, thin stems, and overlapping regions. This stage therefore transforms the standardized initialization into a dense representation. The optimized Gaussian representation
G
G
supports differentiable rendering, whereas the extracted dense point cloud
P
d
P_{d}
serves as the discrete geometric carrier for 2D-to-3D semantic transfer, scale recovery, leaf instance separation, and local trait computation.

3DGS provides an efficient explicit representation for converting sparse initialization into a continuous 3D scene representation
(
Kerbl et al., 2023
)
. However, standard photometric 3DGS may produce aliasing, boundary dilation, and floating artifacts around thin plant structures. Since plant phenotyping requires both smooth leaf surfaces and well-preserved high-frequency structures such as leaf edges, leaf tips, and thin stems, we adopt Mip-Splatting as the geometry-constrained 3DGS densification method
(
Yu et al., 2024
;
Ojo et al., 2024
)
in the main pipeline. Its 3D smoothing and 2D footprint control improve geometric stability under scale variation while maintaining efficient rendering.

Given the camera parameter set
Θ
=
{
(
K
i
,
T
i
)
}
i
=
1
N
\Theta=\{(K_{i},T_{i})\}_{i=1}^{N}
and the initial sparse point cloud
P
s
=
{
p
j
}
j
=
1
M
P_{s}=\{p_{j}\}_{j=1}^{M}
, Gaussian primitives
g
j
=
(
μ
j
,
Σ
j
,
α
j
,
c
j
)
g_{j}=(\mu_{j},\Sigma_{j},\alpha_{j},c_{j})
are initialized from
P
s
P_{s}
, where
μ
j
\mu_{j}
,
Σ
j
\Sigma_{j}
,
α
j
\alpha_{j}
, and
c
j
c_{j}
denote the center, covariance, opacity, and color of the
j
j
-th Gaussian. The representation is then optimized under multi-view photometric constraints:

G
(
0
)
=
Init
(
P
s
)
,
G
=
arg
​
min
G
∑
i
=
1
N
ℒ
photo
(
R
(
G
;
K
i
,
T
i
)
,
I
i
)
,
G^{(0)}=\operatorname{Init}(P_{s}),\qquad G=\operatorname*{arg\,min}_{G}\sum_{i=1}^{N}\mathcal{L}_{\mathrm{photo}}\!\left(R(G;K_{i},T_{i}),I_{i}\right),

where
R
⁡
(
⋅
)
R(\cdot)
denotes differentiable rendering and
ℒ
photo
\mathcal{L}_{\mathrm{photo}}
denotes the photometric reconstruction loss.

To stabilize Gaussian scales and image-space projections, Mip-Splatting constrains both the 3D covariance and the projected 2D footprint:

Σ
~
j
=
Σ
j
+
σ
3
​
D
2
​
I
3
,
Σ
~
i
​
j
2
​
D
=
J
i
​
j
​
Σ
~
j
​
J
i
​
j
⊤
+
σ
2
​
D
2
​
I
2
,
\widetilde{\Sigma}_{j}=\Sigma_{j}+\sigma_{3D}^{2}I_{3},\qquad\widetilde{\Sigma}^{2D}_{ij}=J_{ij}\widetilde{\Sigma}_{j}J_{ij}^{\top}+\sigma_{2D}^{2}I_{2},

where
σ
3
​
D
\sigma_{3D}
and
σ
2
​
D
\sigma_{2D}
control the minimum smoothing scale in 3D space and image space, respectively, and
J
i
​
j
J_{ij}
is the projection Jacobian of Gaussian
g
j
g_{j}
in view
i
i
. The 3D term suppresses degenerate or excessively sharp primitives, while the 2D term reduces aliasing and boundary dilation in projection.

During optimization, Gaussians with large image-space gradients are densified to improve local coverage around undersampled leaf edges, leaf tips, and thin stems, whereas low-opacity, unstable, or isolated primitives are pruned to suppress floating artifacts. After optimization, a dense point cloud is extracted from the Gaussian representation:

P
d
=
S
⁡
(
G
)
.
P_{d}=S(G).

In practice,
S
⁡
(
⋅
)
S(\cdot)
extracts valid Gaussian centers after opacity-based filtering and removal of isolated primitives, producing a discrete point cloud with higher spatial coverage and more continuous boundary structures than the original sparse initialization.

The resulting
G
G
and
P
d
P_{d}
provide complementary outputs for the subsequent pipeline:
G
G
supplies the rendering function needed for view-based operations, while
P
d
P_{d}
provides the dense geometric support for semantic back-projection, scale recovery, leaf separation, and phenotypic measurement. This 3DGS-centered reconstruction-to-measurement workflow is summarized in Fig.
3
.

2.4
Few-view reconstruction via iterative view synthesis and enhancement

Few-view reconstruction remains a major bottleneck for high-throughput plant phenotyping
(
Long et al., 2022
)
. Field, handheld, or robotic acquisition may provide only limited views and overlap, causing SfM initialization failures when feature matching becomes unreliable
(
Schonberger and Frahm, 2016
)
. Neural rendering can also degrade under sparse supervision, producing blurred textures, fragmented leaf geometry, and floating artifacts
(
Niemeyer et al., 2022
)
. In plant scenes, thin leaves, repeated texture, self-occlusion, and open organ boundaries further amplify these errors, which may propagate to semantic transfer, scale recovery, leaf instance separation, and phenotypic measurement.

In this study, few-view mainly refers to highly sparse reconstruction with no more than 10 input views, while larger view budgets were evaluated by uniformly subsampling the full closed-loop sequence to characterize the transition toward dense-view performance. The experiment therefore tests both usability under extremely limited input and view-threshold behavior as observations increase. This module was used only for the few-view reconstruction experiments and ablation settings with novel-view augmentation; full-view reconstruction used the original sampled frames without iterative supplementation unless otherwise specified.

Existing few-view methods generally rely on geometric regularization or generative priors. Regularization-based approaches such as SparseNeuS and FreeNeRF stabilize sparse-view optimization
(
Long et al., 2022
;
Yang et al., 2023
)
, but may oversmooth leaf edges, tips, and fine stems. Generative priors can improve visual quality but may introduce cross-view or geometric inconsistencies. Therefore, generated views are used only as additional image observations that provide supplementary photometric constraints, not as geometric ground truth.

Built upon the 3DFM initialization and Mip-Splatting-based densification developed in the previous stages, we introduce an iterative scheme of novel-view synthesis, view refinement, input update, and re-reconstruction, as illustrated in
Fig.
3
a
. Novel-view poses are sampled between existing viewpoints along the closed-loop trajectory, prioritizing large angular gaps. The current Gaussian representation renders intermediate views at these poses, and Difix3D+
(
Wu et al., 2025
)
refines the rendered images. Difix3D+ is used in its public pretrained form, without further training or fine-tuning, to reduce residual rendering artifacts while using the original sparse views as structural references for geometric consistency.

Figure 3
:
Pipeline for few-view 3D plant reconstruction, segmentation, and phenotyping.

The workflow starts from sparse multi-view inputs, performs 3DFM-based initialization and Mip-Splatting-based densification, supplements effective observations through Difix3D+-assisted view refinement, and then converts the reconstructed geometry into organ-level traits through 2D-to-3D semantic transfer, scale recovery, leaf instance separation, and phenotypic measurement.

Let
I
(
t
)
I^{(t)}
denote the input view set at iteration
t
t
,
G
(
t
)
G^{(t)}
the current Gaussian representation, and
Θ
~
(
t
)
\widetilde{\Theta}^{(t)}
the sampled novel-view poses. The rendering and refinement process is written as

I
~
(
t
)
=
R
⁡
(
G
(
t
)
,
Θ
~
(
t
)
)
,
I
^
(
t
)
=
F
Difix
​
(
I
~
(
t
)
,
I
(
t
)
)
,
\widetilde{I}^{(t)}=R\!\left(G^{(t)},\widetilde{\Theta}^{(t)}\right),\qquad\widehat{I}^{(t)}=F_{\mathrm{Difix}}\!\left(\widetilde{I}^{(t)};I^{(t)}\right),

where
R
⁡
(
⋅
)
R(\cdot)
denotes Gaussian rendering and
I
(
t
)
I^{(t)}
provides the current input views as structural references. The refined views are then merged into the input set and fed back to initialization and densification:

I
(
t
+
1
)
=
I
(
t
)
∪
I
^
(
t
)
,
(
Θ
(
t
+
1
)
,
P
s
(
t
+
1
)
)
=
F
init
​
(
I
(
t
+
1
)
)
,
G
(
t
+
1
)
=
F
dens
​
(
Θ
(
t
+
1
)
,
P
s
(
t
+
1
)
)
.
\begin{gathered}I^{(t+1)}=I^{(t)}\cup\widehat{I}^{(t)},\\
(\Theta^{(t+1)},P_{s}^{(t+1)})=F_{\mathrm{init}}\!\left(I^{(t+1)}\right),\\
G^{(t+1)}=F_{\mathrm{dens}}\!\left(\Theta^{(t+1)},P_{s}^{(t+1)}\right).\end{gathered}

Here,
F
init
F_{\mathrm{init}}
denotes 3DFM inference followed by the bridge conversion described above, and
F
dens
F_{\mathrm{dens}}
denotes the Mip-Splatting-based densification stage. Since the 3DFM front end accepts variable numbers of input views, the refined novel views can be incorporated into the next reconstruction round.

In practice, novel-view number and location are determined by input sparsity and angular gaps along the acquisition loop, with priority given to weakly observed regions. Iterative rendering, refinement, and re-reconstruction expand sparse inputs into a more informative observation set, improving local structural continuity and organ-boundary completeness before subsequent 2D-to-3D semantic transfer and phenotypic measurement.

2.5
2D-to-3D semantic transfer

Direct semantic segmentation on plant point clouds remains challenging due to irregular geometry, thin structures, severe self-occlusion, and the limited availability of large-scale annotated 3D datasets
(
Schunck et al., 2021
)
. To address these issues, we adopt a projection-based 2D-to-3D semantic transfer strategy, which transfers semantic information from image space to 3D space through explicit pixel-to-point correspondence and multi-view label fusion
(
Imabuchi and Kawabata, 2024
)
, as shown in
Fig.
3
b
.

Given the dense point cloud
P
d
=
{
p
i
}
i
=
1
M
P_{d}=\{p_{i}\}_{i=1}^{M}
, where
p
i
∈
ℝ
3
p_{i}\in\mathbb{R}^{3}
, its geometric center and spatial scale are computed as

p
¯
=
1
M
​
∑
i
=
1
M
p
i
,
r
=
max
i
⁡
‖
p
i
−
p
¯
‖
.
\bar{p}=\frac{1}{M}\sum_{i=1}^{M}p_{i},\qquad r=\max_{i}\|p_{i}-\bar{p}\|.

Based on this normalized geometry, a fixed set of virtual cameras is placed on a sphere with radius proportional to
r
r
. Camera positions are generated by Fibonacci sampling to obtain approximately uniform angular coverage. To reduce interference from non-target structures such as pots, views with strong pot dominance or poor plant visibility are avoided, and the remaining viewpoints are biased toward the upper hemisphere of the plant canopy.

For each virtual viewpoint, the point cloud is transformed into the camera coordinate system and projected onto the image plane using a perspective projection model:

u
=
f
x
​
X
Z
+
c
x
,
v
=
f
y
​
Y
Z
+
c
y
.
u=f_{x}\frac{X}{Z}+c_{x},\qquad v=f_{y}\frac{Y}{Z}+c_{y}.

A depth-based visibility constraint is applied so that only the closest visible point contributes to each pixel. To improve the continuity of rendered images, especially around thin organs, a local splatting strategy is used, allowing each point to influence a small pixel neighborhood. When multiple points contribute to the same pixel, the closest visible point is recorded in the pixel-to-point index map. This index map explicitly stores the correspondence between rendered pixels and 3D points, enabling direct label back-projection without relying on post-hoc nearest-neighbor search
(
Imabuchi and Kawabata, 2024
)
.

Semantic segmentation is then performed on the rendered images using a SAM-based 2D segmentation module
(
Kirillov et al., 2023
)
.Compared with directly operating on sparse and irregular point clouds, mature 2D segmentation models provide stronger semantic priors for separating plant organs
(
Chen et al., 2018
;
Zhang et al., 2026
)
. The resulting 2D labels are transferred back to
P
d
P_{d}
through the pixel-to-point index map, assigning semantic predictions to the corresponding 3D points.

Since each 3D point may be observed from multiple virtual viewpoints, it can receive multiple semantic predictions. These predictions are aggregated across views, and the final 3D label is determined by majority voting:

L
i
=
arg
​
max
ℓ
∑
k
(
L
i
(
k
)
=
ℓ
)
,
L_{i}=\operatorname*{arg\,max}_{\ell}\sum_{k}\mathbf{1}\!\left(L_{i}^{(k)}=\ell\right),

where
L
i
(
k
)
L_{i}^{(k)}
denotes the label assigned to point
p
i
p_{i}
from the
k
k
-th rendered view, and
ℓ
\ell
denotes the semantic class. This multi-view fusion integrates complementary observations from different directions and reduces the influence of occlusion, local rendering noise, and single-view segmentation errors
(
Dai and Niessner, 2018
)
.

The output of this stage is a labeled dense point cloud
(
P
d
,
L
3
​
D
)
(P_{d},L_{3D})
, where
L
3
​
D
=
{
L
i
}
i
=
1
M
L_{3D}=\{L_{i}\}_{i=1}^{M}
. This semantic point cloud provides the direct input for subsequent scale recovery, leaf instance separation, and organ-level phenotypic measurement.

2.6
Scale recovery and leaf instance separation

The labeled dense point cloud
(
P
d
,
L
3
​
D
)
(P_{d},L_{3D})
obtained from 2D-to-3D semantic transfer already carries semantic labels for each point, but it still cannot be used directly for absolute phenotypic measurement. Two issues remain: the reconstruction-scale point cloud lacks physical metric scale, and the semantic leaf class is still a single point set rather than a set of measurable leaf instances. This section therefore performs metric scale recovery for sequences with a measured in-scene reference and separates semantic leaf points into individual leaf instances, corresponding to
Fig.
3
c
.

For sequences used for absolute trait evaluation, the known pot geometry is used as an in-scene metric reference. Conventional scale recovery often relies on external calibration objects, such as checkerboards or reference spheres
(
Harandi et al., 2023
)
. These markers increase acquisition complexity and are also prone to occlusion by the plant canopy. In contrast, the pot diameter was measured before imaging and used for scale recovery when the pot rim was reliably visible. Sequences without a reliable metric reference were used for reconstruction-scale analysis rather than absolute physical trait measurement.

We first fit the ground plane with RANSAC and align the vertical direction to the
Z
Z
axis, so that the fitted ground plane defines the
X
​
O
​
Y
XOY
plane. The pot-bottom center is then taken as the coordinate origin. The pot region is isolated from the lower part of the reconstructed scene using geometric constraints, and its outer-boundary candidate points are extracted by angular binning in polar coordinates. Circle fitting is then applied to the outer rim of the pot. Let
D
real
D_{\mathrm{real}}
denote the measured pot diameter and
D
rec
D_{\mathrm{rec}}
the reconstructed diameter estimated from the pot rim. The scale factor and scaled point cloud are defined as

D
rec
=
F
pot
​
(
P
d
)
,
s
^
=
D
real
D
rec
,
P
~
d
=
{
s
^
​
(
p
i
−
o
)
∣
p
i
∈
P
d
}
,
D_{\mathrm{rec}}=F_{\mathrm{pot}}(P_{d}),\qquad\hat{s}=\frac{D_{\mathrm{real}}}{D_{\mathrm{rec}}},\qquad\widetilde{P}_{d}=\{\hat{s}(p_{i}-o)\mid p_{i}\in P_{d}\},

where
o
o
denotes the pot-bottom center in the reconstructed coordinate system. After this transformation,
P
~
d
\widetilde{P}_{d}
is represented in a unified physical coordinate system with the
Z
Z
axis as the vertical direction, which provides the reference for plant height and leaf inclination measurement.

Based on the scaled dense point cloud and the 3D semantic labels, the points classified as leaf are extracted as

P
cand
=
{
p
~
i
∈
P
~
d
∣
L
i
=
leaf
}
,
P_{\mathrm{cand}}=\{\tilde{p}_{i}\in\widetilde{P}_{d}\mid L_{i}=\mathrm{leaf}\},

where
L
i
L_{i}
denotes the semantic label of point
p
i
p_{i}
. Because leaves are often connected through petioles, stems, or local contact regions, direct Euclidean clustering or region growing tends to produce severe under-segmentation
(
Zarei et al., 2024
)
. We therefore adopt a geometry-driven hierarchical separation strategy
(
Ma et al., 2023
)
and write leaf instance separation as

P
leaf
=
{
P
leaf
q
}
q
=
1
Q
=
F
recover
​
(
F
conn
​
(
F
erode
​
(
P
cand
)
)
)
.
P_{\mathrm{leaf}}=\{P_{\mathrm{leaf}}^{q}\}_{q=1}^{Q}=F_{\mathrm{recover}}\!\bigl(F_{\mathrm{conn}}(F_{\mathrm{erode}}(P_{\mathrm{cand}}))\bigr).

Here,
F
erode
F_{\mathrm{erode}}
denotes voxel-based iterative erosion, which removes sparse outliers and weak connecting structures caused by petioles, stems, or local leaf contact;
F
conn
F_{\mathrm{conn}}
denotes connectivity analysis, which partitions the eroded points into leaf-core connected components and cuts residual low-density bridges using local density cues; and
F
recover
F_{\mathrm{recover}}
denotes boundary recovery and point reassignment, which assigns previously removed edge and boundary points back to the nearest compatible leaf core.

Through these steps, the labeled dense point cloud is converted into a scaled plant point cloud and a set of metric leaf instances. The scaled point cloud provides the basis for plant-height estimation, whereas the separated leaf instances provide the direct input for leaf area and leaf inclination angle measurement in the subsequent phenotypic computation stage.

2.7
Phenotypic measurement

After scale recovery and leaf instance separation, phenotypic traits are extracted from the scaled plant point cloud and separated leaf instances under a unified physical coordinate system. Let
P
leaf
=
{
P
leaf
q
}
q
=
1
Q
P_{\mathrm{leaf}}=\{P_{\mathrm{leaf}}^{q}\}_{q=1}^{Q}
denote the separated leaf instances, where
P
leaf
q
P_{\mathrm{leaf}}^{q}
is the point cloud of the
q
q
-th leaf. For each instance, we reconstruct a 3D mesh and compute leaf area and leaf inclination angle from surface geometry. Plant height is obtained from the scaled dense point cloud as the vertical distance between the pot-bottom reference plane and the highest plant point, and is used as an auxiliary whole-plant trait. The main quantitative evaluation focuses on leaf area and leaf inclination angle, because they directly reflect organ-level geometric completeness and posture recovery; this final trait-extraction stage is shown in
Fig.
3
d
.

For leaf area estimation, we reconstruct an open mesh for each leaf using the Ball-Pivoting Algorithm (BPA)
(
Bernardini et al., 1999
)
. Compared with closed-surface methods such as Poisson surface reconstruction, BPA is more suitable for thin, open leaf surfaces and reduces overestimation caused by artificial boundary closure. Because leaf phenotyping concerns the one-sided area of a leaf blade, area is computed from the cleaned single-surface mesh after removing duplicated or opposite-side artifacts. Let
T
q
T_{q}
denote the triangles in the reconstructed mesh of the
q
q
-th leaf, where
S
i
S_{i}
and
n
→
i
=
(
n
i
,
x
,
n
i
,
y
,
n
i
,
z
)
\vec{n}_{i}=(n_{i,x},n_{i,y},n_{i,z})
represent the area and normal vector of the
i
i
-th triangle, respectively. Isolated triangles, locally duplicated triangles, and triangles with abnormal area or inconsistent normals are removed. The remaining valid triangle set is denoted as
T
~
q
⊆
T
q
\tilde{T}_{q}\subseteq T_{q}
. The one-sided area of the
q
q
-th leaf is then defined as

A
q
=
∑
i
∈
T
~
q
S
i
.
A_{q}=\sum_{i\in\tilde{T}_{q}}S_{i}.

Leaf inclination angle is defined consistently with the manual measurement as the acute angle between the fitted mean leaf plane and the horizontal plane, equivalently computed as the acute angle between the mean leaf surface normal and the global vertical direction
(
Itakura and Hosoi, 2019
)
. To reduce local noise and mesh irregularity, triangle normals are oriented consistently within each leaf mesh and then averaged with area weights:

n
→
q
=
∑
i
∈
T
~
q
S
i
​
n
→
i
‖
∑
i
∈
T
~
q
S
i
​
n
→
i
‖
.
\vec{n}_{q}=\frac{\sum_{i\in\tilde{T}_{q}}S_{i}\,\vec{n}_{i}}{\left\|\sum_{i\in\tilde{T}_{q}}S_{i}\,\vec{n}_{i}\right\|}.

Let
v
→
=
(
0
,
0
,
1
)
\vec{v}=(0,0,1)
denote the global vertical vector established during scale recovery. The leaf inclination angle of the
q
q
-th leaf is then computed as

θ
q
=
arccos
⁡
(
|
n
→
q
⋅
v
→
|
)
.
\theta_{q}=\arccos\left(\left|\vec{n}_{q}\cdot\vec{v}\right|\right).

The absolute value removes normal-direction ambiguity, ensuring a consistent physical reference across plants and leaf instances.

The outputs are the auxiliary whole-plant height and paired leaf-level traits
{
A
q
,
θ
q
}
q
=
1
Q
\{A_{q},\theta_{q}\}_{q=1}^{Q}
, which are compared with manual measurements in the quantitative evaluation.

3
Results

We systematically evaluate the key components of the proposed framework, including fast reconstruction, few-view recovery, segmentation strategy, and phenotypic parameter extraction. Using experiments on multiple crops and diverse scenes, we assess the method’s overall performance with respect to reconstruction efficiency, geometric quality, organ-level segmentation, and quantitative measurement accuracy.

3.1
Overall framework

To evaluate visual geometry 3D Foundation Models as the front-end geometric foundation for plant 3D phenotyping, this study extends the evaluation scope from reconstruction quality alone to the full phenotyping workflow. The core question is whether low-cost image inputs can achieve second-level stable initialization, and then successfully go through dense reconstruction, semantic perception, scale recovery, and leaf instance separation, ultimately producing comparable organ-level phenotypic traits. The Results section follows the data flow of front-end geometric recovery, few-view reconstruction, 3D semantic perception, terminal phenotypic measurement, providing a staged evaluation of the proposed pipeline.

As illustrated in Fig.
1
, the framework takes smartphone video frames or sampled multi-view images as input, denoted as
I
=
{
I
i
}
i
=
1
N
I=\{I_{i}\}_{i=1}^{N}
. In the first stage, the 3DFM-based front end, together with bridge conversion, recovers standardized camera parameters and initial sparse geometry:

(
Θ
,
P
s
)
\displaystyle(\Theta,P_{s})

=
F
init
​
(
I
)
=
F
convert
​
(
F
3
​
D
​
F
​
M
m
​
(
I
)
)
,
\displaystyle=F_{\mathrm{init}}(I)=F_{\mathrm{convert}}\!\left(F_{\mathrm{3DFM}}^{m}(I)\right),

Θ
\displaystyle\Theta

=
{
(
K
i
,
T
i
)
}
i
=
1
N
.
\displaystyle=\{(K_{i},T_{i})\}_{i=1}^{N}.

Here,
Θ
\Theta
denotes the set of camera intrinsics and poses,
P
s
P_{s}
denotes the initial sparse point cloud, and
F
init
F_{\mathrm{init}}
represents 3DFM inference followed by bridge conversion.

In the second stage, geometry-constrained 3DGS converts the sparse initialization into a continuous 3D representation:

G
=
F
dens
​
(
Θ
,
P
s
)
,
P
d
=
S
⁡
(
G
)
.
\begin{gathered}G=F_{\mathrm{dens}}(\Theta,P_{s}),\\
P_{d}=S(G).\end{gathered}

Here,
G
G
denotes the Gaussian representation,
P
d
P_{d}
denotes the dense point cloud extracted from the optimized representation. This stage provides novel-view rendering capacity, more complete leaf boundaries, fine stems, canopy structures, serving as the geometric basis for downstream segmentation, measurement. Under sparse-view input, the framework supplements effective observations through view synthesis, image enhancement, re-reconstruction, reducing the dependence of plant 3D reconstruction on dense acquisition.

In the third stage, the dense 3D result enters the semantic perception process. A 2D-to-3D semantic transfer strategy projects multi-view 2D segmentation results back to 3D space through explicit pixel–point correspondence, then obtains stable 3D semantic labels through multi-view fusion:

L
3
​
D
=
F
seg
​
(
P
d
,
𝒱
)
.
L_{3D}=F_{\mathrm{seg}}(P_{d};\mathcal{V}).

Here,
𝒱
\mathcal{V}
denotes the virtual viewpoints used for projection, 2D segmentation, back-projection, and multi-view fusion.

After semantic point clouds are obtained, pot geometry is used for metric scale recovery, followed by leaf instance separation to split the leaf class into individual leaf objects:

(
s
^
,
P
~
d
)
=
F
scale
​
(
P
d
,
D
real
)
,
P
leaf
=
{
P
leaf
q
}
q
=
1
Q
=
F
inst
​
(
P
~
d
,
L
3
​
D
)
.
\begin{gathered}(\hat{s},\widetilde{P}_{d})=F_{\mathrm{scale}}(P_{d},D_{\mathrm{real}}),\\
P_{\mathrm{leaf}}=\{P_{\mathrm{leaf}}^{q}\}_{q=1}^{Q}=F_{\mathrm{inst}}(\widetilde{P}_{d},L_{3D}).\end{gathered}

The separated metric leaf instances are then used to compute phenotypic traits, including leaf area, leaf inclination angle, and plant height:

τ
=
F
trait
​
(
P
~
d
,
P
leaf
)
=
{
H
,
{
(
A
q
,
θ
q
)
}
q
=
1
Q
}
.
\tau=F_{\mathrm{trait}}(\widetilde{P}_{d},P_{\mathrm{leaf}})=\left\{H,\{(A_{q},\theta_{q})\}_{q=1}^{Q}\right\}.

Through this design, the pipeline of reconstruction, segmentation and measurement forms a continuous data-transformation chain aimed at rapid generation of measurable 3D plant structures from low-cost images.

Based on this framework, the following results address four questions. First, whether 3DFM-based initialization can reduce the conventional minute-level front end to the second level across multi-crop, multi-scene data, while achieving reconstruction quality comparable to COLMAP. Second, whether iterative view synthesis and enhancement can lower the view threshold required for usable reconstruction under limited input views. Third, whether 2D-to-3D semantic transfer can provide stable organ perception on complex plant point clouds. Fourth, whether reconstructed structures, after scale recovery, leaf instance separation, can support quantitative phenotypic measurements such as leaf area, leaf inclination angle. This validation sequence follows the workflow in Fig.
1
, evaluating the proposed plant 3D phenotyping route in terms of front-end efficiency, geometric stability, semantic perception capacity, and terminal phenotypic accuracy.

3.2
3DFM-based initialization reduces the plant reconstruction front end from minutes to seconds

Table 1
:
Comparison of reconstruction performance using VGGT, COLMAP, and
π
3
\pi^{3}
on different plant datasets.
Time is reported in minutes (m) or seconds (s).

Method

Tobacco_Outdoor_01

Maize_Outdoor_01

Maize_Field_01

Tobacco_Outdoor_02

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

COLMAP

0.893

30.372

0.130

6.189 m

0.869

29.491

0.161

5.141 m

0.877

28.874

0.134

8.131 m

0.901

29.784

0.128

6.862 m

VGGT

0.875

29.347

0.165

1.865 s

0.833

28.989

0.189

1.619 s

0.855

27.235

0.179

1.354 s

0.865

28.913

0.167

1.427 s

π
3
\pi^{3}

0.889

30.218

0.133

1.602 s

0.870

29.487

0.167

1.881 s

0.869

28.128

0.141

0.987 s

0.891

29.297

0.134

1.390 s

Method

Wheat_Outdoor_01

Maize_Indoor_01

Tobacco_Field_01

Bamboo_Indoor_01

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

COLMAP

0.872

28.219

0.157

4.178 m

0.889

29.214

0.138

5.982 m

0.902

29.845

0.129

7.341 m

0.881

28.967

0.143

6.127 m

VGGT

0.853

27.213

0.183

1.585 s

0.861

28.337

0.174

1.534 s

0.875

28.914

0.168

1.693 s

0.854

27.842

0.181

1.382 s

π
3
\pi^{3}

0.874

28.179

0.153

1.464 s

0.884

29.102

0.142

1.487 s

0.895

29.613

0.135

1.446 s

0.876

28.721

0.152

1.764 s

Method

Tobacco_Outdoor_03

Tobacco_Outdoor_04

Tobacco_Outdoor_05

Tobacco_Outdoor_06

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

COLMAP

0.893

29.514

0.137

8.492 m

0.878

28.391

0.149

4.913 m

0.904

29.914

0.126

6.871 m

0.890

29.304

0.141

5.613 m

VGGT

0.862

28.347

0.176

1.919 s

0.851

27.144

0.185

1.487 s

0.876

28.841

0.171

1.693 s

0.863

28.267

0.177

1.447 s

π
3
\pi^{3}

0.887

29.289

0.141

1.631 s

0.873

28.256

0.155

1.603 s

0.899

29.712

0.132

1.592 s

0.886

29.187

0.148

1.729 s

Method

Wheat_Indoor_01

Wheat_Field_01

Wheat_Field_02

Soybean_Indoor_01

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

COLMAP

0.892

29.612

0.139

7.284 m

0.883

28.934

0.148

4.721 m

0.897

29.728

0.131

6.512 m

0.905

29.981

0.124

8.193 m

VGGT

0.865

28.457

0.176

1.512 s

0.857

27.823

0.182

1.403 s

0.868

28.593

0.171

1.557 s

0.879

28.942

0.165

1.684 s

π
3
\pi^{3}

0.889

29.471

0.147

1.683 s

0.879

28.713

0.155

1.516 s

0.892

29.501

0.138

1.746 s

0.901

29.718

0.129

1.533 s

Method

Soybean_Indoor_02

Wheat_Outdoor_02

Rapeseed_Indoor_01

Legume_Indoor_01

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

COLMAP

0.886

29.148

0.144

5.923 m

0.893

29.487

0.138

7.612 m

0.882

28.963

0.146

4.983 m

0.898

29.631

0.133

6.388 m

VGGT

0.859

28.013

0.178

1.312 s

0.867

28.415

0.171

1.406 s

0.855

27.942

0.184

1.475 s

0.871

28.592

0.169

1.621 s

π
3
\pi^{3}

0.881

28.934

0.151

1.728 s

0.889

29.315

0.145

1.557 s

0.879

28.798

0.152

1.629 s

0.894

29.508

0.140

1.507 s

Method

Pea_Indoor_01

Pea_Indoor_02

Pea_Indoor_03

Pea_Indoor_04

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

COLMAP

0.887

29.184

0.142

5.712 m

0.892

29.415

0.138

7.981 m

0.899

29.784

0.130

9.182 m

0.884

28.742

0.147

4.312 m

VGGT

0.862

28.193

0.176

1.422 s

0.869

28.447

0.172

1.774 s

0.872

28.619

0.167

1.422 s

0.859

27.894

0.181

1.305 s

π
3
\pi^{3}

0.884

29.041

0.148

1.638 s

0.889

29.294

0.144

1.582 s

0.895

29.593

0.137

1.693 s

0.881

28.593

0.153

1.445 s

Method

Broccoli_Indoor_01

Wheat_Indoor_02

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

SSIM
↑
\uparrow

PSNR
↑
\uparrow

LPIPS
↓
\downarrow

Time
↓
\downarrow

COLMAP

0.891

29.621

0.139

6.583 m

0.903

29.947

0.128

8.342 m

VGGT

0.866

28.417

0.173

1.654 s

0.876

28.913

0.169

1.723 s

π
3
\pi^{3}

0.887

29.478

0.146

1.578 s

0.900

29.815

0.135

1.568 s

Front-end initialization represents the first bottleneck in the proposed phenotyping pipeline. Subsequent 3DGS densification, few-view enhancement, 2D-to-3D semantic back-projection, scale recovery, and leaf-level measurement all require stable camera parameters and initial geometry. To test whether 3DFMs can serve as the geometric entry point for plant 3D phenotyping, we compared COLMAP, VGGT, and
π
3
\pi^{3}
on
26
26
plants covering diverse crops, morphologies, and acquisition scenes. Table
1
reports initialization efficiency and initialization-driven reconstruction quality using SSIM, PSNR, LPIPS, and front-end runtime. Runtime was restricted to initialization: feature extraction, matching, pose estimation, and sparse reconstruction for COLMAP; model inference plus bridge conversion to standardized camera parameters and sparse geometry for VGGT and
π
3
\pi^{3}
. We excluded downstream densification, scale recovery, and phenotypic measurement from the runtime evaluation.

The main gain was efficiency. COLMAP required an average front-end runtime of
6.52
​
min
6.52~\mathrm{min}
, with a median of
6.45
​
min
6.45~\mathrm{min}
. VGGT and
π
3
\pi^{3}
required only
1.55
​
s
1.55~\mathrm{s}
and
1.58
​
s
1.58~\mathrm{s}
on average, with medians of
1.52
​
s
1.52~\mathrm{s}
and
1.60
​
s
1.60~\mathrm{s}
, respectively. Based on mean runtime, VGGT and
π
3
\pi^{3}
achieved approximately
252.6
×
252.6\times
and
247.3
×
247.3\times
speedups over COLMAP. Thus, 3DFM-based initialization compressed the dominant front-end cost from minutes to seconds, with consistent behavior across all samples.

This acceleration remained useful because the quality loss was limited, especially for
π
3
\pi^{3}
. COLMAP achieved the highest average SSIM, PSNR, and LPIPS values of
0.890
0.890
,
29.387
​
dB
29.387~\mathrm{dB}
, and
0.138
0.138
.
π
3
\pi^{3}
reached
0.886
0.886
,
29.191
​
dB
29.191~\mathrm{dB}
, and
0.144
0.144
, remaining close to COLMAP. VGGT reached
0.863
0.863
,
28.333
​
dB
28.333~\mathrm{dB}
, and
0.175
0.175
, showing a larger drop. Relative to COLMAP,
π
3
\pi^{3}
reduced mean SSIM by only
0.004
0.004
, reduced PSNR by
0.196
​
dB
0.196~\mathrm{dB}
, and increased LPIPS by
0.006
0.006
; VGGT reduced mean SSIM by
0.027
0.027
, reduced PSNR by
1.055
​
dB
1.055~\mathrm{dB}
, and increased LPIPS by
0.036
0.036
.
π
3
\pi^{3}
therefore provided an approximately
250
×
250\times
acceleration with only a minor reconstruction-quality cost.

The comparison between
π
3
\pi^{3}
and VGGT determined the default initializer for subsequent experiments. Both models enabled second-level feed-forward visual geometry inference, yet
π
3
\pi^{3}
achieved higher SSIM and PSNR than VGGT on all
26
26
samples, together with consistently lower LPIPS. This sample-wide advantage indicates more stable initialization for complex plant canopies, thin leaf boundaries, and local occlusions. Since subsequent 3DGS optimization, semantic back-projection, and leaf-level measurement depend on camera–point geometric consistency,
π
3
\pi^{3}
provides a more suitable geometric starting point for the full pipeline.

Together, Table
1
establishes 3DFM-based initialization as a practical alternative to COLMAP-style front ends in the proposed framework. COLMAP defines a high-quality conventional baseline, VGGT demonstrates the feasibility of unified feed-forward visual geometry inference for plant reconstruction, and
π
3
\pi^{3}
provides the best quality–efficiency balance. With the front-end cost reduced from minutes to seconds, the next question is whether the 3DFM-driven pipeline can remain usable when the number of input views is further reduced.

3.3
Comparison of performance in few-view reconstruction

Figure 4
:
Effect of Input View Count on Reconstruction and Phenotyping Performance.
PSNR, SSIM, and LPIPS are reported for COLMAP and the proposed
π
3
\pi^{3}
+ novel view synthesis pipeline under different numbers of input views. The proposed method achieves usable reconstruction at much lower view counts and approaches the dense-view performance ceiling more rapidly.

After front-end initialization was compressed from the minute level to the second level, the bottleneck of high-throughput plant 3D phenotyping shifted further toward image acquisition. In field inspection, robotic platforms, and handheld rapid capture, view number and view overlap are often constrained by time, trajectory, and occlusion. To evaluate the usability of the proposed workflow under low-view input, we compared COLMAP versus the
π
3
\pi^{3}
+ novel-view synthesis pipeline across different input view counts, using reconstruction quality and phenotypic measurement accuracy as evaluation targets. The key question is whether stable 3D geometry can still be recovered under sparse views, then support leaf area and leaf inclination estimation.

As shown in Fig.
4
, the reconstruction-quality curves show a clear view-threshold effect. In the low-view range, COLMAP failed to produce valid reconstructions, with PSNR and SSIM close to zero and LPIPS close to the failure ceiling, indicating that matching-based SfM cannot reliably establish camera poses and initial geometry under low-overlap input. The
π
3
\pi^{3}
+ novel-view synthesis pipeline recovered rapidly in the same range, with PSNR, SSIM, and LPIPS entering the usable range earlier. This difference indicates that the primary bottleneck in few-view reconstruction is whether initialization can be established;
π
3
\pi^{3}
provides a stable front-end geometry, while novel-view synthesis and refinement supplement missing observations, thereby reducing the number of views required for usable reconstruction.

As the view count increased, COLMAP showed a sharp improvement at approximately
30
30
–
45
45
views, indicating that conventional SfM requires sufficient view overlap to enter a stable reconstruction regime. The
π
3
\pi^{3}
+ novel-view synthesis pipeline had already approached a plateau before this threshold. After more than
50
50
views, the PSNR, SSIM, and LPIPS curves of the two methods gradually converged. This trend shows that the main benefit of the proposed method lies in the low-view regime: its core value is shifting the usable reconstruction threshold forward, rather than raising the performance ceiling under dense-view input.

The phenotypic measurement curves further show that the few-view advantage propagates to terminal traits. The
R
2
R^{2}
values of leaf area and leaf inclination angle increased with view count. COLMAP could not produce valid phenotypic estimates in the reconstruction-failure range, whereas the
π
3
\pi^{3}
+ novel-view synthesis pipeline reached higher accuracy with fewer views. The increase in leaf-area
R
2
R^{2}
reflects more stable scale recovery, leaf boundaries, and instance separation. The increase in inclination-angle
R
2
R^{2}
indicates more reliable leaf-surface geometry, vertical reference, and normal estimation. Overall, this experiment shows that the proposed workflow reduces the view threshold required for both 3D reconstruction and organ-level phenotypic measurement, making it suitable for low-cost rapid acquisition.

3.4
2D-to-3D semantic transfer achieves more consistent plant organ segmentation than direct 3D approaches

Following reconstruction enhancement and point cloud densification, we compared two segmentation paradigms for reconstructed plant point clouds: direct 3D semantic segmentation and projection-based 2D-to-3D semantic transfer. The 3D baselines included PSegNet, TPointNet++, and PointTransformerV3, whereas the 2D-to-3D branch used the proposed pipeline of multi-view rendering, label back-projection, and multi-view fusion. For training the 3D baselines, we constructed a cross-crop dataset by integrating SoybeanMVS
(
Sun et al., 2023
)
, MaizeField3D
(
Kimara et al., 2026
)
, and syau-single-maize
(
Yang et al., 2024
)
, and split it into training and validation sets at an 8:2 ratio to support multi-crop segmentation with a single model. We then evaluated all trained models on our reconstructed plant dataset.

Figure 5
:
Qualitative comparison of plant organ segmentation results across methods.

Qualitative results of PSegNet, TPointNet++, PointTransformerV3, and the proposed method on soybean, maize, and sesame samples, with GT shown for reference.

As shown in Fig.
5
, Direct 3D segmentation showed clear limitations when applied to plant point cloud scenarios. Boundaries between leaves and stems were frequently blurred by thin organs, severe occlusion, uneven point density, and local holes, resulting in fragmentation, incorrect merging, and class confusion. Cross-crop structural variation further reduced stability, since maize exhibits a stronger vertical organization, whereas soybean and sesame contain denser branching and more entangled organs.

The projection-based 2D-to-3D strategy produced more consistent results across crops. The densified point cloud was first rendered into multiple views, and a pixel-to-point index map was constructed to establish explicit correspondence between image space and 3D space. Semantic predictions were then obtained in 2D, back-projected to the point cloud, and fused across views. This design preserved leaf boundaries more effectively, improved robustness under occlusion, and reduced the ambiguity introduced by post-hoc nearest-neighbor reprojection.

Table 2
:
Quantitative comparison of plant organ segmentation performance on different crops.

Crop

Method

OA
↑
\uparrow

mIoU
↑
\uparrow

Leaf IoU
↑
\uparrow

Stem IoU
↑
\uparrow

Soybean1

PSegNet

0.7028

0.4656

0.6673

0.2639

TPointNet++

0.7561

0.5144

0.7275

0.3014

PointTransformerV3

0.7676

0.5174

0.7429

0.2920

Ours

0.9660

0.9185

0.9543

0.8827

Soybean2

PSegNet

0.7524

0.4164

0.7465

0.0863

TPointNet++

0.7364

0.3914

0.7327

0.0501

PointTransformerV3

0.6337

0.3504

0.6219

0.0790

Ours

0.9362

0.8216

0.9237

0.7195

Maize1

PSegNet

0.7809

0.3962

0.7803

0.0121

TPointNet++

0.7880

0.4891

0.7759

0.2023

PointTransformerV3

0.9075

0.5073

0.9065

0.1081

Ours

0.9962

0.9810

0.9957

0.9662

Maize2

PSegNet

0.9222

0.5943

0.9199

0.2688

TPointNet++

0.7821

0.4411

0.7763

0.1060

PointTransformerV3

0.9062

0.5162

0.9049

0.1275

Ours

0.9956

0.9816

0.9950

0.9682

Sesame

PSegNet

0.8482

0.5899

0.8351

0.3446

TPointNet++

0.7847

0.4999

0.7699

0.2299

PointTransformerV3

0.7664

0.4854

0.7498

0.2209

Ours

0.9742

0.9096

0.9698

0.8495

The quantitative comparison in Table
2
further supports these observations. The proposed method consistently achieved the best OA, mIoU, Leaf IoU, and Stem IoU across all evaluated crops. Overall, the advantage of 2D-to-3D segmentation lies in its better match to reconstructed plant data. It reduces dependence on large-scale point-level annotations, transfers more reliably across crops, and handles thin, overlapping structures more effectively through multi-view fusion and explicit geometric correspondence. We therefore adopt 2D-to-3D semantic transfer as the core strategy for subsequent leaf instance separation and phenotypic analysis.

3.5
Leaf area and inclination angle are reliably estimated from reconstructed 3D plant structures

After reconstruction and segmentation, the final test of the pipeline is whether reconstructed and segmented 3D structures can be converted into reliable organ-level traits. We evaluated leaf area and leaf inclination angle as terminal indicators. Leaf area depends on metric scale, leaf boundary preservation, instance separation, and mesh-based area computation. Leaf inclination angle depends on leaf-surface continuity, the vertical reference, and single-leaf normal estimation. These two traits evaluate scale-related and posture-related phenotypes, respectively.

Across five representative soybean, maize, and sesame plants, the estimated traits remained highly consistent with manual measurements (Fig.
6
). Leaf-area regression reached
R
2
=
0.9362
R^{2}=0.9362
–
0.9438
0.9438
, and leaf-inclination regression reached
R
2
=
0.9307
R^{2}=0.9307
–
0.9455
0.9455
, indicating that the pipeline preserved inter-leaf differences in both size and orientation. At the plant level, the
Estimated
/
GT
\mathrm{Estimated}/\mathrm{GT}
ratio of total leaf area ranged from
0.9514
0.9514
to
1.0629
1.0629
, corresponding to an approximately
−
4.86
%
-4.86\%
to
+
6.29
%
+6.29\%
deviation. The ratios remained close to
1.0
1.0
, suggesting limited global scale distortion. The mean absolute error of leaf inclination angle was approximately
2.04
∘
2.04^{\circ}
, indicating stable coordinate normalization and normal estimation for posture traits.

Crop-specific patterns reflected structural differences in measurement difficulty. Maize leaves are elongated and axis-dominant, favoring boundary preservation and normal estimation; maize1 reached
R
2
=
0.9438
R^{2}=0.9438
for leaf area, and maize2 reached
R
2
=
0.9449
R^{2}=0.9449
for inclination angle. Soybean showed more overlap, petiole connection, and local occlusion; Soybean1 had the lowest leaf-area regression, with
R
2
=
0.9362
R^{2}=0.9362
, reflecting greater boundary-recovery difficulty in compound leaves. Sesame maintained
R
2
=
0.9363
R^{2}=0.9363
for leaf area and
R
2
=
0.9307
R^{2}=0.9307
for inclination angle, supporting transfer across different leaf forms and canopy structures.

Figure 6
:
Phenotypic measurement accuracy across crops and growth stages.

Top: cross-crop agreement between estimated and manual leaf area and inclination angle. Bottom: stage-wise leaf-area Estimated/GT ratios and signed leaf-inclination deviations, with shaded reference ranges shown for area ratios and inclination deviations, respectively.

Stage-wise results further showed measurement stability. Across four growth stages, the
Estimated
/
GT
\mathrm{Estimated}/\mathrm{GT}
ratio for leaf area fluctuated around
1.0
1.0
, indicating no cumulative area error with plant growth and leaf expansion. Inclination-angle deviations mostly stayed within the
±
5
∘
\pm 5^{\circ}
reference range, indicating stable vertical-reference recovery and leaf-normal estimation across stages. This result extends the validation from single-time-point cross-crop measurement to continuous growth-stage phenotyping, suggesting that the proposed pipeline has the potential to support dynamic phenotypic recording. Together, the cross-crop regression, stage-wise stability, and low angular error show that the pipeline preserves the geometric, semantic, and metric consistency required for organ-level measurement, enabling low-cost image-based reconstruction to be converted into comparable 3D plant phenotypic traits.

3.6
All three design choices improve performance, with sparse initialization contributing the strongest gain

Table
3
presents ablation results showing consistent gains from all three design choices, with the strongest effect arising from sparse initialization. Under matched dense modeling and augmentation settings,
π
3
\pi^{3}
improves PSNR, area
R
2
R^{2}
, and angle
R
2
R^{2}
over COL while also reducing the overall runtime. Mip consistently outperforms standard 3DGS in paired configurations, indicating better preservation of thin leaf boundaries and local structural continuity. Novel-view augmentation brings smaller but systematic gains, mainly by compensating for sparse observations and refining local geometry.

Table 3
:
Ablation study of the proposed framework with module-wise runtime.

No.

Sparse

init.

Dense

model

Novel-view

aug.

PSNR
↑
\uparrow

Area

R
2
↑
R^{2}\uparrow

Angle

R
2
↑
R^{2}\uparrow

Init.

time

Dense

time

(min)

Aug.

time

(min)

Total

time

(min)
↓
\downarrow

1

COL

3DGS

×
\times

27.98

0.898

0.910

6.4 min

30.9

0.0

37.3

2

COL

3DGS

✓
\checkmark

28.24

0.905

0.917

6.4 min

30.9

1.2

38.5

3

COL

Mip

×
\times

28.56

0.916

0.925

6.4 min

31.5

0.0

37.9

4

COL

Mip

✓
\checkmark

28.95

0.924

0.934

6.4 min

31.5

1.2

39.1

5

π
3
\pi^{3}

3DGS

×
\times

29.04

0.919

0.926

1.6 s

30.8

0.0

30.8

6

π
3
\pi^{3}

3DGS

✓
\checkmark

29.30

0.926

0.932

1.6 s

30.8

1.2

32.0

7

π
3
\pi^{3}

Mip

×
\times

29.58

0.935

0.940

1.6 s

31.4

0.0

31.4

8

π
3
\pi^{3}

Mip

✓
\checkmark

29.86

0.941

0.946

1.6 s

31.4

1.2

32.6

Note:
COL denotes COLMAP initialization.
π
3
\pi^{3}
initialization is reported in seconds, whereas dense reconstruction, novel-view augmentation, and total runtime are reported in minutes. Dense time is computed as total time minus sparse initialization and novel-view augmentation time, and is rounded to one decimal place.

At the configuration level, No.8 (
π
3
\pi^{3}
+ Mip + novel-view augmentation) achieves the best PSNR, area
R
2
R^{2}
, and angle
R
2
R^{2}
, demonstrating that the framework’s advantage stems from the coupling of stable initialization, measurement-oriented densification, and sparse-view enhancement rather than any single module alone. Notably, No.5 gives the shortest runtime, whereas No.8 gives the highest accuracy, indicating that the added cost of novel-view augmentation is limited but sufficient to further release the potential of the
π
3
\pi^{3}
–Mip combination.

4
Discussion

4.1
From 3DFM-based initialization to measurable phenotyping

This study positions visual geometry 3DFMs as the front-end geometric foundation for plant 3D phenotyping and tests whether this change propagates to organ-level measurement. COLMAP-style initialization relies on feature extraction, matching, and incremental pose recovery, making it vulnerable to repetitive texture, occlusion, and thin organs. In contrast,
π
3
\pi^{3}
reduces the average front-end runtime from
6.52
​
min
6.52~\mathrm{min}
to
1.58
​
s
1.58~\mathrm{s}
while maintaining reconstruction quality close to COLMAP, providing an efficient geometric entry point for 3DGS densification, few-view enhancement, and 3D semantic perception.

This front-end shift also relaxes acquisition constraints. Combined with novel-view synthesis,
π
3
\pi^{3}
lowers the usable reconstruction threshold and recovers stable geometry from fewer views. The pipeline links faster initialization, synthesized observations, boundary-preserving 3DGS, 2D-to-3D semantic transfer, scale recovery, and leaf instance separation, converting reconstructed geometry into metric leaf-level objects. Thus, measurable traits require not only visual fidelity, but also geometric continuity, semantic consistency, and physical scale.

The phenotypic results verify this system-level propagation. Across representative crops, the plant-level
Estimated
​
-
​
to
​
-
​
GT
\mathrm{Estimated}\text{-}\mathrm{to}\text{-}\mathrm{GT}
ratio of total leaf area ranges from
0.9514
0.9514
to
1.0629
1.0629
, and the mean absolute error of leaf inclination angle is approximately
2.04
∘
2.04^{\circ}
. The near-unity area ratios and low angular error indicate limited scale bias, stable vertical reference recovery, leaf-surface reconstruction, and normal estimation. Thus, 3DFM-based initialization establishes a verifiable geometric starting point from low-cost images to reconstruction, perception, and organ-level phenotypic extraction.

4.2
Current limitations and applicability boundaries

The current framework is most reliable for single-plant, close-range, closed-loop acquisition or structurally separated scenes, where viewpoint continuity and target visibility support coherent geometry across initialization, densification, semantic transfer, and measurement. Dense canopies, dynamic disturbances, cluttered backgrounds, and persistent occlusion remain challenging for camera recovery, local completion, cross-view semantic consistency, and leaf instance separation. Therefore, the present results mainly support low-cost, close-range, isolated-plant phenotyping, with further validation required in field-scale canopies.

Metric scale recovery is another boundary. Absolute traits such as leaf area and plant height require a reliable scale anchor. This study uses pot geometry as an in-scene reference, reducing calibration cost in potted or controlled acquisition, but this is only one anchoring strategy. In field deployment, scale may come from platform pose, camera height, ground-plane constraints, row or plant spacing, or reference markers. Without such an anchor, the workflow supports relative structural comparison and temporal change analysis rather than absolute traits.

Local organ separation remains challenging. Leaf-area and leaf-inclination errors are affected by edge completeness, petiole connections, leaf contact, occlusion, and curled structures, which complicate 2D-to-3D label transfer, instance separation, and mesh reconstruction. Although current results show stable overall accuracy, complex canopies still require leaf-level error analysis, failure visualization, and larger-scale validation. Future work should address marker-free scale recovery, robust organ separation, and higher-order phenotypes such as curvature, torsion, stem topology, and temporal growth dynamics. At this stage, the framework should be regarded as validated for low-cost, close-range, organ-level 3D phenotyping, with dense field canopies and marker-free deployment requiring further validation.

4.3
Implications for low-cost and cross-crop plant phenotyping

Within the applicability boundaries discussed above, this study shows that smartphone images can support measurable organ-level 3D plant phenotyping. Low-cost acquisition, 3DFM-based front-end initialization, and few-view reconstruction reduce hardware, time, and data barriers, while geometry-constrained 3DGS, 2D-to-3D semantic transfer, scale recovery, and leaf instance separation convert image-based reconstruction into measurement-ready objects with physical scale and organ boundaries. Thus, low-cost 3D phenotyping depends on coordinated reductions in acquisition, reconstruction, and measurement conversion.

To our knowledge, this study is among the first to systematically introduce visual geometry 3DFMs into plant 3D phenotyping and validate them as a front-end geometric foundation linking image acquisition, 3D reconstruction, and organ-level measurement. In this framework, 3DFMs provide a unified entry point for camera recovery, initial structure formation, few-view completion, and downstream 3D optimization. Accordingly, plant 3D phenotyping evaluation should extend beyond rendering quality to semantic consistency, metric scale, organ-instance quality, and final trait accuracy.

The cross-crop dataset and manually annotated ground truth support system-level evaluation beyond reconstruction quality, covering few-view stability, 3D perception, and organ-level measurement. Overall, this study establishes a 3DFM-driven, measurement-oriented route for low-cost cross-crop 3D plant phenotyping, connecting second-level initialization, few-view reconstruction, semantic perception, dataset-supported validation, and organ-level trait extraction into a complete evidence chain. The framework remains mainly suited to close-range, single-plant, or structurally separated acquisition, and future extensions are needed for complex field canopies, marker-free scale recovery, and higher-order dynamic traits.

5
Conclusions

This study presents a 3DFM-driven workflow for low-cost, cross-crop 3D plant phenotyping. By replacing COLMAP-style matching-based initialization with feed-forward 3D foundation model inference, the proposed pipeline compresses the reconstruction front end from minutes to seconds while maintaining reconstruction quality close to conventional methods. Combined with Mip-Splatting-based densification, few-view supplementation, 2D-to-3D semantic transfer, metric scale recovery, and leaf instance separation, the framework converts low-cost image inputs into measurable organ-level 3D traits. Experiments across diverse crops and acquisition conditions demonstrate that the proposed route supports rapid reconstruction, robust segmentation, and reliable estimation of leaf area and inclination angle. Future work should further extend the method to dense field canopies, marker-free scale recovery, stronger organ separation under occlusion, and multi-temporal dynamic phenotyping.

CRediT authorship contribution statement

Wei Zhou and Hanyue Jia contributed equally to this work. All authors jointly conceived and designed the study. Hanyue Jia and Wei Zhou developed the methodology and performed the experiments. Wenbo Zhou and Yanan Li contributed to data analysis and experimental validation. Hanyue Jia and Wei Zhou prepared the initial manuscript. Hao Lu and Tingting Wu supervised the study, reviewed the manuscript, and provided critical revisions. All authors read and approved the final manuscript.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Funding

This work was supported by the National Natural Science Foundation of China under Grant No. 62576146.

Acknowledgements

The authors gratefully acknowledge the financial support from the National Natural Science Foundation of China.

References

Bernardini
et al.
(1999)

F. Bernardini, J. Mittleman, H. Rushmeier, C. Silva, and G. Taubin

The ball-pivoting algorithm for surface reconstruction
.

IEEE Transactions on Visualization and Computer Graphics

5
(
4
),
pp. 349–359
.

External Links:
Document

Cited by:
§2.7
.

Chen
et al.
(2018)

L. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam

Encoder-decoder with atrous separable convolution for semantic image segmentation
.

In
Proceedings of the European Conference on Computer Vision (ECCV)
,

pp. 801–818
.

Cited by:
§2.5
.

Dai and Niessner (2018)

A. Dai and M. Niessner

3DMV: joint 3d-multi-view prediction for 3d semantic scene segmentation
.

In
Proceedings of the European Conference on Computer Vision (ECCV)
,

pp. 452–468
.

Cited by:
§2.5
.

Deng
et al.
(2026)

W. Deng, T. Wu, Z. Ni, Y. Liu, H. Jia, and Q. Ling

Hyperspectral imaging meets 3d gaussian splatting: a novel approach beyond 3d plant morphology
.

Plant Phenomics

8
(
2
),
pp. 100198
.

External Links:
Document

Cited by:
§1
.

El Banani
et al.
(2024)

M. El Banani, A. Raj, K. Maninis, A. Kar, Y. Li, M. Rubinstein, D. Sun, L. Guibas, J. Johnson, and V. Jampani

Probing the 3d awareness of visual foundation models
.

In
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
,

pp. 21795–21806
.

Cited by:
§1
.

Furbank and Tester (2011)

R. T. Furbank and M. Tester

Phenomics—technologies to relieve the phenotyping bottleneck
.

Trends in Plant Science

16
(
12
),
pp. 635–644
.

External Links:
Document

Cited by:
§1
.

Furukawa and Ponce (2010)

Y. Furukawa and J. Ponce

Accurate, dense, and robust multiview stereopsis
.

IEEE Transactions on Pattern Analysis and Machine Intelligence

32
(
8
),
pp. 1362–1376
.

External Links:
Document

Cited by:
§1
.

Harandi
et al.
(2023)

N. Harandi, B. Vandenberghe, J. Vankerschaver, S. Depuydt, and A. Van Messem

How to make sense of 3d representations for plant phenotyping: a compendium of processing and analysis techniques
.

Plant Methods

19
(
1
),
pp. 60
.

External Links:
Document

Cited by:
§1
,

§2.2
,

§2.6
.

Houle
et al.
(2010)

D. Houle, D. R. Govindaraju, and S. Omholt

Phenomics: the next challenge
.

Nature Reviews Genetics

11
(
12
),
pp. 855–866
.

External Links:
Document

Cited by:
§1
.

Hu
et al.
(2025)

T. Hu, P. Shen, Y. Zhang,
et al.

OpenPheno: an open-access, user-friendly, and smartphone-based software platform for instant plant phenotyping
.

Plant Methods

21
,
pp. 76
.

Cited by:
§1
,

§2.1
.

Imabuchi and Kawabata (2024)

R. Imabuchi and K. Kawabata

Discrimination of plant structures in 3d point cloud through back-projection of labels derived from 2d semantic segmentation
.

Journal of Robotics and Mechatronics

36
(
1
),
pp. 63–70
.

External Links:
Document

Cited by:
§2.5
,

§2.5
.

Itakura and Hosoi (2019)

K. Itakura and F. Hosoi

Estimation of leaf inclination angle in three-dimensional plant images obtained from lidar
.

Remote Sensing

11
(
3
),
pp. 344
.

External Links:
Document

Cited by:
§2.1
,

§2.7
.

Kerbl
et al.
(2023)

B. Kerbl, G. Kopanas, T. Leimkuhler, and G. Drettakis

3D gaussian splatting for real-time radiance field rendering
.

ACM Transactions on Graphics

42
(
4
),
pp. 139:1–139:14
.

External Links:
Document

Cited by:
§1
,

§1
,

§2.3
.

Kimara
et al.
(2026)

E. Kimara, M. Hadadi, J. Godbersen, A. Balu, T. Z. Jubery, Y. Li, A. Krishnamurthy, P. S. Schnable, and B. Ganapathysubramanian

MaizeField3D: a curated 3d point cloud and procedural model dataset of field-grown maize from a diversity panel
.

Plant Phenomics

8
(
1
),
pp. 100108
.

External Links:
Document

Cited by:
§3.4
.

Kirillov
et al.
(2023)

A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W. Lo, P. Dollár, and R. Girshick

Segment anything
.

In
Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
,

pp. 4015–4026
.

External Links:
Document

Cited by:
§2.5
.

Leroy
et al.
(2024)

V. Leroy, Y. Cabon, and J. Revaud

Grounding image matching in 3d with mast3r
.

In
Computer Vision – ECCV 2024
,

pp. 73–90
.

External Links:
Document

Cited by:
§1
,

§2.2
.

Li
et al.
(2025)

J. Li, X. Qi, Z. Fu, Y. Wang, G. Xiong, D. Zeng, X. Wang, X. Liu, S. Teng, F. Hosoi,
et al.

A survey on 3d reconstruction techniques in plant phenotyping: from classical methods to neural radiance fields, 3d gaussian splatting, and beyond
.

Plant Phenomics

7
,
pp. 100137
.

External Links:
Document

Cited by:
§1
.

Liu
et al.
(2023)

R. Liu, R. Wu, B. Van Hoorick, P. Tokmakov, S. Zakharov, and C. Vondrick

Zero-1-to-3: zero-shot one image to 3d object
.

Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
,
pp. 9298–9309
.

Cited by:
§1
.

Long
et al.
(2022)

X. Long, C. Lin, P. Wang, T. Komura, and W. Wang

SparseNeuS: fast generalizable neural surface reconstruction from sparse views
.

In
Computer Vision – ECCV 2022
,

pp. 210–227
.

External Links:
Document

Cited by:
§1
,

§2.4
,

§2.4
.

Luo
et al.
(2023)

L. Luo, X. Jiang, Y. Yang, E. R. A. Samy, M. Lefsrud, V. Hoyos-Villegas, and S. Sun

Eff-3dpseg: 3d organ-level plant shoot segmentation using annotation-efficient point clouds
.

Plant Phenomics
.

Cited by:
§1
,

§1
.

Ma
et al.
(2023)

Z. Ma, R. Du, J. Xie, D. Sun, H. Fang, L. Jiang, and H. Cen

Phenotyping silique morphology in oilseed rape using skeletonization with hierarchical segmentation
.

Plant Phenomics

5
,
pp. 0027
.

External Links:
Document

Cited by:
§2.6
.

Mildenhall
et al.
(2020)

B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng

NeRF: representing scenes as neural radiance fields for view synthesis
.

In
Computer Vision – ECCV 2020
,

pp. 405–421
.

External Links:
Document

Cited by:
§1
.

Niemeyer
et al.
(2022)

M. Niemeyer, J. T. Barron, B. Mildenhall, M. S. M. Sajjadi, A. Geiger, and N. Radwan

RegNeRF: regularizing neural radiance fields for view synthesis from sparse inputs
.

In
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
,

pp. 5480–5490
.

Cited by:
§2.4
.

Ojo
et al.
(2024)

T. Ojo, T. La, A. Morton, and I. Stavness

Splanting: 3d plant capture with gaussian splatting
.

ACM SIGGRAPH Asia 2024 Technical Communications
,
pp. 1–4
.

External Links:
Document

Cited by:
§1
,

§2.1
,

§2.3
.

Okura (2022)

F. Okura

3D modeling and reconstruction of plants and trees: a cross-cutting review across computer graphics, vision, and plant phenotyping
.

Breeding Science

72
(
1
),
pp. 31–45
.

External Links:
Document

Cited by:
§1
,

§1
,

§1
,

§2.2
.

Schonberger and Frahm (2016)

J. L. Schonberger and J. Frahm

Structure-from-motion revisited
.

In
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
,

pp. 4104–4113
.

External Links:
Document

Cited by:
§1
,

§2.2
,

§2.4
.

Schunck
et al.
(2021)

D. Schunck, F. Magistri, R. A. Rosu, J. Cornelissen, N. Chebrolu, S. Paulus, P. Lottes, and C. Stachniss

Pheno4D: a spatio-temporal dataset of maize and tomato plant point clouds for phenotyping and advanced plant analysis
.

PLoS ONE

16
(
8
),
pp. e0256340
.

External Links:
Document

Cited by:
§2.5
.

Shen
et al.
(2025)

P. Shen, J. Xueyao, W. Deng, H. Jia, and T. Wu

PlantGaussian: exploring 3d gaussian splatting for cross-time, cross-scene, and realistic 3d plant visualization and beyond
.

The Crop Journal

13
(
2
),
pp. 607–618
.

Cited by:
§1
.

Song
et al.
(2025)

W. Song, H. Huang, Y. Sun, F. Qu, J. Zhang, L. Fang, Y. Hao, and C. Peng

IPENS: interactive unsupervised framework for rapid plant phenotyping extraction via nerf-sam2 fusion
.

Plant Phenomics
.

Cited by:
§1
,

§1
.

Sun
et al.
(2023)

Y. Sun, Z. Zhang, K. Sun, S. Li, J. Yu, L. Miao, Z. Zhang, Y. Li, H. Zhao, Z. Hu, D. Xin, Q. Chen, and R. Zhu

Soybean-mvs: annotated three-dimensional model dataset of whole growth period soybeans for 3d plant organ segmentation
.

Agriculture

13
(
7
),
pp. 1321
.

External Links:
Document

Cited by:
§3.4
.

Tao
et al.
(2022)

H. Tao, S. Xu, C. Miao, H. Long, S. Jin, and Q. Guo

Proximal and remote sensing in plant phenomics: 20 years of progress, challenges, and perspectives
.

Plant Communications

3
(
6
),
pp. 100344
.

External Links:
Document

Cited by:
§1
.

Wang
et al.
(2025)

H. Wang
et al.

AMB3R: accurate feed-forward metric-scale 3d reconstruction
.

arXiv preprint arXiv:2511.20343
.

Cited by:
§1
.

Wang
et al.
(2025a)

J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny

VGGT: visual geometry grounded transformer
.

Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
.

Cited by:
§1
,

§1
,

§1
,

§2.2
.

Wang
et al.
(2024)

S. Wang, V. Leroy, Y. Cabon, B. Chidlovskii, and J. Revaud

DUSt3R: geometric 3d vision made easy
.

Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
,
pp. 20697–20709
.

Cited by:
§1
,

§2.2
.

Wang
et al.
(2025b)

Y. Wang, J. Zhou, H. Zhu, W. Chang, Y. Zhou, Z. Li, J. Chen, J. Pang, C. Shen, and T. He

π
3
\pi^{3}
: permutation-equivariant visual geometry learning
.

arXiv preprint arXiv:2507.13347
.

Cited by:
§1
,

§1
,

§1
,

§2.2
.

Wu
et al.
(2025)

J. Z. Wu, Y. Zhang, H. Turki, X. Ren, J. Gao, M. Z. Shou, S. Fidler, Z. Gojcic, and H. Ling

DIFIX3D+: improving 3d reconstructions with single-step diffusion models
.

In
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
,

pp. 26024–26035
.

Cited by:
§2.4
.

Xu
et al.
(2022)

Y. Xu, X. Zhang, H. Li, H. Zheng, J. Zhang, M. S. Olsen, R. K. Varshney, B. M. Prasanna, and Q. Qian

Smart breeding driven by big data, artificial intelligence, and integrated genomic-enviromic prediction
.

Molecular Plant

15
(
11
),
pp. 1664–1695
.

External Links:
Document

Cited by:
§1
.

Yang
et al.
(2023)

J. Yang, M. Pavone, and Y. Wang

FreeNeRF: improving few-shot neural rendering with free frequency regularization
.

In
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
,

pp. 8254–8263
.

Cited by:
§2.4
.

Yang
et al.
(2026)

N. Yang
et al.

Three-dimensional phenotyping: technological advances and recent applications in crop research
.

Plant Communications
.

Cited by:
§1
.

Yang
et al.
(2020)

W. Yang, H. Feng, X. Zhang, J. Zhang, J. H. Doonan, W. D. Batchelor, L. Xiong, and J. Yan

Crop phenomics and high-throughput phenotyping: past decades, current challenges, and future perspectives
.

Molecular Plant

13
(
2
),
pp. 187–214
.

External Links:
Document

Cited by:
§1
.

Yang
et al.
(2025)

X. Yang
et al.

PlantSegNeRF: a few-shot, cross-species method for plant instance point cloud reconstruction from multi-view images
.

Artificial Intelligence in Agriculture
.

Cited by:
§1
,

§1
.

Yang
et al.
(2024)

X. Yang, T. Miao, X. Tian, D. Wang, J. Zhao, L. Lin, C. Zhu, T. Yang, and T. Xu

Maize stem–leaf segmentation framework based on deformable point clouds
.

ISPRS Journal of Photogrammetry and Remote Sensing

211
,
pp. 49–66
.

External Links:
Document

Cited by:
§3.4
.

Yu
et al.
(2024)

Z. Yu, A. Chen, M. Tancik, D. Charatan, A. Geiger, R. Ng, J. T. Barron, B. Mildenhall, P. Hedman, G. Kopanas, T. Leimkuhler, G. Drettakis, L. Fan, N. Snavely, G. Riegler, S. Liu, Q. Wang, and A. Dai

Mip-splatting: alias-free 3d gaussian splatting
.

In
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
,

pp. 19447–19456
.

Cited by:
§1
,

§1
,

§2.3
.

Zarei
et al.
(2024)

A. Zarei, B. Li, J. C. Schnable, E. Lyons, D. Pauli, B. Benes, and K. Barnard

PlantSegNet: 3d point cloud instance segmentation of nearby plant organs with identical semantics
.

Computers and Electronics in Agriculture

221
,
pp. 108922
.

External Links:
Document

Cited by:
§2.6
.

Zhang
et al.
(2025a)

D. Zhang, J. Gajardo, T. Medic, I. Katircioglu, M. Boss, N. Kirchgessner, A. Walter, and L. Roth

Wheat3DGS: in-field 3d reconstruction, instance segmentation and phenotyping of wheat heads with gaussian splatting
.

In
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops
,

pp. 5399–5409
.

Cited by:
§1
,

§1
.

Zhang
et al.
(2026)

J. Zhang, S. Cao, B. Xu, Y. Li, W. Jia, T. Wu, H. Lu, W. Hu, and Z. Han

DepthCropSeg++: scaling a crop segmentation foundation model with depth-labeled data
.

arXiv preprint arXiv:2601.12366
.

Cited by:
§2.5
.

Zhang
et al.
(2025b)

S. Zhang, J. Wang, Y. Xu, N. Xue, C. Rupprecht, X. Zhou, Y. Shen, and G. Wetzstein

FLARE: feed-forward geometry, appearance and camera estimation from uncalibrated sparse views
.

In
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
,

pp. 21936–21947
.

Cited by:
§1
.

Zheng
et al.
(2026)

Y. Zheng, L. Gao, J. Zhang, L. Miao, H. Zhu, X. Yang, D. Zhang, Z. Han, X. Li, and W. Zhu

Hi magicring, tell me where i am: toward affordable, physically reliable 3d plant phenotyping with mobilepheno3d
.

aBIOTECH

7
(
3
),
pp. 100045
.

External Links:
Document

Cited by:
§1
.

Experimental support, please

view the build logs

for errors. Generated by

L

A

T

E

xml

.

Instructions for reporting errors

We are continuing to improve HTML versions of papers, and your feedback helps enhance accessibility and mobile
support. To report errors in the HTML that will help us improve conversion and rendering, choose any of the
methods listed below:

Click the "Report Issue"
(
)
button, located in the page header.

Tip:
You can select the relevant text first, to include it in your report.

Our team has already identified
the following issues
. We appreciate your time reviewing and reporting rendering errors we
may not have found yet. Your efforts will help us improve the HTML versions for all readers, because disability
should not be a barrier to accessing research. Thank you for your continued support in championing open access for
all.

Have a free development cycle? Help support accessibility at arXiv! Our collaborators at LaTeXML maintain a
list of packages that need conversion
, and welcome
developer contributions
.

We gratefully acknowledge support from
our
major funders
,

member institutions
,
,
and all contributors.

About

·

Help

·

Contact

·

Subscribe

·

Copyright

·

Privacy

·

Accessibility

·

Operational Status
(opens in new tab)

Major funding support from
</reference>

<statements>
1. This report proposes a structured modeling, analysis, and design framework for such phenotyping systems, focusing on state-space models, observers, optimal and model predictive control (MPC), and stochastic estimation, together with complementary methods from computer vision and machine learning.
2. Finally, the report outlines practical design patterns and architectures for an agricultural engineering researcher to build field- or lab-scale grain phenotyping platforms with explicit control-theoretic performance guarantees.
3. Foundation-model-based pipelines integrate 3D scene priors with organ-level semantic segmentation, enabling organ separation, metric scale recovery, and trait extraction in seconds rather than minutes. These systems are particularly promising for high-throughput grain phenotyping, where thousands of seeds or panicles must be processed.
4. Because 3DGS and 3D foundation models dramatically reduce reconstruction latency—from minutes to seconds—closed-loop scheduling based on reconstructed geometry becomes feasible. This opens the door to real-time adaptation in field phenotyping, for example steering UAV flights around plots to target regions where traits like panicle density show high uncertainty.
5. Wheat3DGS and foundation-model pipelines use 3DGS plus multi-view segmentation to extract hundreds of wheat heads and measure length, width, and volume automatically.
6. 3D foundation models provide pre-trained geometric priors that streamline reconstruction across many crops. They replace traditional matching-based SfM initialization with feed-forward inference, then refine geometry via 3DGS.
7. Rapid reconstruction from smartphone images in seconds.
8. Choose reconstruction algorithms (SfM–MVS, NeRF, 3DGS, 3DFMs) based on environment and scale.
9. Implement closed-loop MPC or feedback controllers for adaptive imaging in dynamic conditions.
10. For an agricultural engineering researcher, adopting explicit state-space models, quality metrics, and optimal control/estimation strategies can transform ad hoc imaging setups into principled measurement systems with quantifiable performance guarantees and clear paths to scaling and field deployment.
</statements>

Begin the assessment now. Output only the JSON list, without any conversational text or explanations.