You will be provided with a research report. The body of the report will contain some citations to references.

Citations in the main text may appear in the following forms:
1. A segment of text + space + number, for example: "Li Qiang constructed a socioeconomic status index (SES) based on income, education, and occupation, dividing society into 7 levels 15"
2. A segment of text + [number], for example: "Li Qiang constructed a socioeconomic status index (SES) based on income, education, and occupation, dividing society into 7 levels[15]"
3. A segment of text + [number†(some line numbers, etc.)], for example: "Li Qiang constructed a socioeconomic status index (SES) based on income, education, and occupation, dividing society into 7 levels[15†L10][5L23][7†summary]"
4. [Citation Source](Citation Link), for example: "According to [ChinaFile: A Guide to Social Class in Modern China](https://www.chinafile.com/reporting-opinion/media/guide-social-class-modern-china)'s classification, Chinese society can be divided into nine strata"

Please identify **all** instances where references are cited in the main text, and extract (fact, ref_idx, url) triplets. When extracting, pay attention to the following:
1. Since these facts will need to be verified later, you may need to look for some context before and after the citation to ensure that the fact is complete and understandable, rather than just a simple phrase or short expression.
2. If a fact cites multiple references, then it should correspond to two triplets: (fact, ref_idx_1, url_1) and (fact, ref_idx_2, url_2).
3. For the third form of citation (i.e., where the citation source and link appear directly in the text), the ref_idx should be uniformly set to 0.
4. If the main text does not specify the exact location of the citation (for example, only the reference list is listed at the end of the article, without specifying the citation point in the text), please return an empty list.

You should return a JSON list format, where each item in the list is a triplet, for example:
[
    {
        "fact": "Text segment from the original document. Note that Chinese quotation marks should use full-width marks. And add a single backslash before the English quotation mark to make it a readable for python json module.",
        "ref_idx": "The index of the cited reference in the reference list for this text segment.",
        "url": "The URL of the cited reference for this text segment (extracted from the reference list at the end of the research report or from the parentheses at the citation point)."
    }
]

Here is the main text of the research report:
## Executive Summary

The published grain, field and reconstruction pipelines justify treating the scanner [1], motion stage or robot [2][4], illumination [4], focus [5], reconstruction software [14] and trait-analysis layer [1] as one controllable dynamic system. The evidence supports three load-bearing conclusions, each developed below from quantified pipeline results [1][2][10][14] and control results [5][7][8][9].

First, active structured-light and laser-triangulation rigs can deliver sub-millimetre grain geometry, with a Reeyee Pro structured-light scanner reporting 0.05 mm single-sided accuracy within 2 s [1][6], while a low-cost line-laser binocular wheat-grain platform reports length, width, thickness and volume RMSEs of 0.10–0.23 mm and 0.65–1.77 mm³ depending on scan speed [2]. Second, passive and neural-rendering routes reduce hardware cost and capture time but transfer error to reconstruction, registration and occlusion: in one wheat study, 3D Gaussian splatting and NeRF produced point clouds with 0.74 mm and 1.43 mm average accuracy against a handheld scanner, while MVS and SfM produced 2.32 mm and 7.23 mm errors [14], and a consumer maize-ear pipeline reached kernel-count R² = 0.921 but only relative volume proxies [10].

Third, modern estimation and predictive control already exist for the platform and imaging channel: loosely coupled EKF fusion maintains pose when GPS or VIO fails [7], hybrid visual servoing adapts unknown camera intrinsics, extrinsics and robot dynamics with Lyapunov convergence [8], nonlinear moving-horizon estimation plus nonlinear model predictive control achieves field tracking below 5 cm under unknown wheel–terrain slip [9], and optical-model sharpness mapping closes autofocus to more than 95% in-focus accuracy within five iterations over a 35× depth-of-field range [5].

The design contribution is therefore not to invent a new sensor, but to couple these methods into a supervisory digital twin [3] that optimizes scan speed [2], viewpoint [13], focus [5], illumination [4] and scheduling [3] against a trait-uncertainty cost, while explicitly budgeting registration, mixed-pixel, laser-penetration, occlusion and benchmark-transfer errors that the literature quantifies only separately [12][13][11][14].

## System boundary and control variables

The design problem begins from a practical bottleneck: manual grain measurement is inefficient, labor-intensive and subjective [1], and conventional 3D acquisition equipment can be complicated and expensive for high-throughput phenotyping [2]. Three-dimensional sensing is more data-intensive than two-dimensional methods but gives more accurate geometry and better tracks movement, growth and yield over time [16]. It also enables non-destructive longitudinal monitoring [12]. The technology is nevertheless constrained by speed, availability, portability, spatial resolution and cost [16], and the complexity of 3D analysis has been identified as a main bottleneck to wider deployment [16].

A useful design taxonomy is active versus passive acquisition [16]. Active methods use controlled structured energy, such as a scanning laser or projected light pattern, with a camera or detector; they can be costly but return highly accurate data [16]. Passive methods rely on ambient light and commodity hardware, but often produce lower-quality data requiring substantial computation [16]. For small cereal grains, binocular stereo, structure-from-motion and space-carving point clouds are relatively sparse, whereas structured light can obtain high-precision point clouds [1]. For breeding-scale work, technologies also differ by scale, from single-plant laboratory settings to miniplots, experimental fields and open fields [12].

The measured dependencies in the cited pipelines—stage speed [2], acquisition geometry [6], focus [5], illumination and camera height [4], and plant motion [13]—support abstracting the controllable system as:

\[
x_{k+1}=f(x_k,u_k,w_k), \quad y_k=h(x_k,v_k), \quad \hat z_k=g(\hat y_k),
\]

where \(x\) is the platform and object state [4][13], \(u\) is the manipulated acquisition vector [2][5][6], \(y\) is the image or depth measurement [1][14], and \(z\) is the extracted trait vector [1][6][10]. The variables are not arbitrary: scan speed is a measured accuracy–throughput variable [2], turntable rotation, scanning angle and stage colour affect structured-light reconstruction [6], focus is a dynamic optical state [5], illumination and camera height are constrained by robot power and crop proximity [4], and plant motion is caused by wind, egomotion and physiological movement [13].

| Design layer | State variables | Manipulated variables | Evidence constraining the design |
| --- | --- | --- | --- |
| Mechanical platform | stage or table position, robot pose, grain orientation, plant motion | scan speed, turntable angle, robot speed, camera height | WG-3D uses a carrier table at 25 mm/s and shows improved accuracy at 5 mm/s with reduced efficiency [2]; PhenoRob-F reaches a maximum speed of 1 m/s and mounts the camera at different heights with approximately 40 cm working distance [4]; wind and plant egomotion move leaves during scanning [13]. |
| Optical channel | defocus distance, sharpness, illumination, background colour | focus adjustment, lamp power or height, exposure, stage colour | an optical defocus model gives a sharpness-to-distance map for autofocus [5]; halogen lamps consume more than 1200 W and over 80% of total system power [4]; black stage colour was selected as optimal in orthogonal experiments [6]; plant reflectance is highest near 550 nm and above 780 nm [13]. |
| Reconstruction | camera intrinsics and extrinsics, registration residual, point density, mixed-pixel score | viewpoint selection, number of images, fiducial use, filtering strategy | SfM resolution depends on image count, viewing angles and CCD resolution [12]; SfM pose failure can be screened by average trajectory error above 1.5 mm [14]; registration needs fiducial networks because ICP is prone to local minima in unstructured crop scenes [13]; mixed pixels can be flagged by deviation from the expected Gaussian return distribution [13]. |
| Phenotypic analysis | trait vector, class probabilities, genotype and environment labels | segmentation thresholds, feature set, machine-learning model, scheduling priority | 25 rice-grain traits were extracted and used with six machine-learning algorithms [1]; 28 wheat-grain 3D traits and 4 ventral-sulcus traits were extracted [6]; maize-ear kernels were segmented and reprojected into 3D trait sets [10]; plane-segmentation and cluster-range thresholds were user-modifiable in grain software [1]. |

## Modeling the imaging–reconstruction–phenotyping chain

The measurement chain has three coupled models: a motion model, an imaging model and a trait-extraction model. The motion model is the robot or stage. In the structured-light grain system, the scanner, robot, fixture, object platform, industrial computer and control unit form an integrated acquisition cell [1]. In WG-3D, grains are placed on a carrier table and scanned at an even speed, with the line-laser binocular camera acquiring 3D spatial information through binocular parallax and stereo matching [2]. In field platforms, pose is estimated from GNSS, visual navigation and wheel control; PhenoRob-F uses RTK positioning with centimetre-level accuracy when stationary and an improved deviation-coupling algorithm for wheel-speed compensation [4].

The imaging model depends on the selected representation. Structured light combines a projector, two cameras and an internal modulated light source, using triangulation and sinusoidal grating images to obtain dense point clouds [1]. The Reeyee Pro scanner used in rice and wheat grain studies achieved 0.05 mm single-sided accuracy within 2 s [1][6], with a point distance of 0.16 mm, spatial resolution of 0.05 mm, scanning area of 210 × 150 mm, working distance of 290–480 mm and maximum scan size of 200 × 200 × 200 mm [6]. WG-3D instead uses a line-laser binocular camera costing only a few hundred dollars and computes disparity from left and right image coordinates [2]. Passive SfM uses sets of RGB images to reconstruct 3D models, but its resolution depends on image count, viewing-angle diversity and camera-chip resolution [12]. Neural-rendering methods add another representation layer: 3DGS and NeRF can generate high-fidelity novel views and point clouds, while MVS and SfM produce larger average errors when compared against a handheld scanner [14].

The trait-extraction model is the measurement map from point clouds or rendered geometry to phenotypes. In the rice structured-light study, length, width, thickness, surface area and volume were calculated by specified grain-point-cloud algorithms after segmentation, and 25 phenotypic traits were extracted, including compactness index [1]. In WG-3D, the pipeline includes point-cloud acquisition, preprocessing, single-grain extraction, 3D model construction and morphological feature calculation; alpha-shape clustering, overhead projection, minimum bounding rectangles and a thickness formula \(T = Z_{\max}-Z_{\text{ref}}\) are used [2]. In the wheat structured-light study, single-grain segmentation used region growth, dimensions were extracted by constructing an oriented bounding box from an axis-aligned bounding box, projected area was reconstructed by a greedy projection algorithm, and ventral-sulcus depth, perimeter and area were computed from slice geometry [6]. In maize ears, the calibrated point cloud is Z-axis aligned via PCA, cylindrically unwrapped to 2D, contrast-enhanced with CLAHE, segmented with zero-shot Cellpose-SAM using a triple-juxtaposed unwrap to resolve the projection seam, and reprojected to per-kernel 3D point sets for traits including kernel count, kernel row number, kernels per row, surface area, volume proxy, packing geometry, hue and aspect ratio [10].

The synthesis is that the “sensor model” is not fixed. It changes with actuator settings and representation. Structured-light accuracy is high but sensitive to stage colour, scanning angle and rotation angle [6]. WG-3D is low-cost and high-throughput but single-sided, so the back of the grain belly is unavailable [2]. Consumer SfM and neural-rendering pipelines are scalable but inherit registration and reconstruction error [14][10]. Therefore, the control design must optimize acquisition parameters against downstream trait uncertainty, not merely against image quality.

## Identification, calibration and observers

The cited systems require identification at three levels: static acquisition calibration [6], recursive platform and optical state estimation [5][7][8][9], and supervisory digital-twin modelling [3].

Static identification is already evidenced by design-of-experiments. A three-level orthogonal experiment over rotation angle, scanning angle and stage colour on 125 grains from five wheat varieties selected 30° rotation, 37° scanning angle and black stage colour as the minimum-trait-error operating point [6]. The same study showed that blindly increasing rotation angle does not significantly improve accuracy and can reduce measurement efficiency [6]. This is an empirical static map from actuator settings to output quality [6].

Recursive identification is needed for pose, camera calibration and focus. A loosely coupled extended Kalman filter fuses IMU, odometer, GPS and visual-inertial odometry; GPS failure is defined by absence of signal and VIO failure by inter-frame distance exceeding a threshold, with failed sensor values replaced by wheel-odometer position, quaternion and covariance [7]. The authors report that the fusion algorithm is more stable and robust than MSCKF-VIO and IMU-odometer fusion, and they test disabled-sensor intervals [7]. Hybrid visual servoing goes further: it uses three adaptive laws to estimate unknown camera intrinsics, extrinsics and robot dynamics, with Lyapunov proof of asymptotic closed-loop convergence [8]. This is a natural template for an uncalibrated grain-scanning rig because camera and plant parameters are identified inside the control loop [8].

For the optical channel, autofocus can be modelled as a state-estimation problem [5]. The defocused image is represented as \(I_d=I_0*G(D_{\text{coc}})+\text{noise}\), where a Gaussian blur kernel encodes defocus [5]. A deep network predicts sharpness from defocused images, and an adaptive adjustment algorithm uses a sharpness-to-distance map to move the focus stage [5]. The method explicitly addresses defocus uncertainty, where the same defocus level occurs at two axial positions, and achieves over 95% in-focus accuracy within five iterations across a 35× depth-of-field range [5]. Performance varies by texture and illumination: printed paper is harder than anodized or PCB samples, and bright conditions can cause saturation-related prediction errors [5].

For terrain and wheel–slip identification, a nonlinear moving-horizon estimator identifies unknown traction parameters online, constrained between 0 and 1, and feeds a learning-based nonlinear model predictive controller [9]. The controller achieves mean Euclidean tracking error of 0.0423 m after staying on track and remains accurate despite 35 missing GNSS points out of 1850 [9]. The authors note that hard-coding could be more accurate if soil conditions were static, implying that the estimator is most valuable under changing terrain [9].

SfM and neural-rendering pipelines also require calibration and convergence checks [12][14]. SfM resolution depends on image count, viewing angles and camera resolution [12]. In one wheat reconstruction study, SfM processes were judged failed when average trajectory error exceeded 1.5 mm, and only 12 of 20 reconstructions met that criterion [14]. A released field dataset includes RGB images with calibrated camera poses and corresponding laser scans [15]. The design implication is that camera pose and reconstruction convergence must be treated as estimated states with quality flags, not as assumed constants [12][14][15].

## Observability and error budget

Observability is a central limitation of single-view 3D grain phenotyping [12][13][16]. A single stationary viewpoint produces unavoidable data gaps and occlusions, especially in dense canopies, and multiple viewpoints are needed to cover the object [13]. Depth maps fail to capture occluded objects [16], while point clouds include depth information that helps work around leaf occlusion [16]. True 3D models from multiple views reduce occlusion and improve resolution and accuracy compared with 2.5D single-view methods [12]. Even so, occlusion remains present because the inner centre of a plant can be hidden by its own leaves [12]. For individual grains, WG-3D explicitly cannot recover the back side of the grain belly from single-sided scanning [2].

Registration is another observability bottleneck [13]. Point clouds from different viewpoints are stored in local coordinate systems and must be transformed into a common global frame; registration errors cause re-occurrence, displacement and fracturing of plant parts and therefore systematic biases in geometrical traits [13]. ICP is prone to local minima in unstructured, repetitive, changing and landmark-free crop scenes, motivating a fiducial network of black-and-white planar targets with at least 100 points per target, at least four visible targets per scan, and targets enclosing a large surface area or volume [13]. Advanced registration may require scene understanding, such as separating stable soil from unstable plant regions and optimizing each differently [13].

The available evidence does not provide one consolidated per-trait variance decomposition, but it supplies separate quantified terms—placement and scanner error [1][6], speed-dependent point-cloud error [2], registration and mixed-pixel error [13], representation error [11][14] and environmental thresholds [3]—that can be assembled into a design budget.

| Error term | Quantified evidence | Design consequence |
| --- | --- | --- |
| Scanner geometry and point density | Reeyee Pro reports 0.05 mm single-sided accuracy in 2 s [1][6]; the rice grain point cloud reached an average minimum point distance of 0.1731 mm [1]; WG-3D reports length RMSE 0.2256 mm at 25 mm/s and 0.1018 mm at 5 mm/s [2]. | Sub-millimetre grain traits require active benchtop sensing and controlled stage speed [1][2][6]. |
| Grain placement and orientation | Horizontal placement produced length, width and thickness errors of 4.55%, 4.05% and 3.82%, while vertical placement produced 2.15%, 0.68% and 1.18% [1]; wheat and corn were generally more stable than rice in vertical placement [1]; WG-3D cannot recover the back of the grain belly [2]. | Fixtures, orientation control and multi-view or projection strategies must be part of the plant model [1][2]. |
| Registration and pose | ICP is prone to local minima in crop scenes [13]; fiducial-network rules are needed for accurate registration [13]; SfM pose convergence can be screened by ATE > 1.5 mm [14]; TLS-to-3DGS alignment was assessed as within 10 mm on average by visual inspection [15]. | The controller should reject or re-plan scans when pose or registration residuals exceed thresholds [13][14]. |
| Mixed pixels and laser–matter interaction | Mixed pixels increase with footprint size, defeat local-neighbourhood outlier removal, and their removal can bias traits such as making leaf size systematically too small [13]; multi-return separation works only for foreground–background distances above 0.75 m [13]; laser penetration gives below-millimetre systematic bias and can increase noise by an order of magnitude [13]; plant reflectance peaks near 550 nm and above 780 nm [13]. | Wavelength selection, mixed-pixel filtering and distance-aware acquisition are required [13]. |
| Reconstruction representation | Volume carving with 3 views gives approximately +10% systematic volume error, falling to about 0.1% at 35 or 36 images [11]; a deep-learning method with 3 views reaches about 2% relative error but L1 loss systematically underestimates volume [11]; 3DGS and NeRF point clouds show 0.74 mm and 1.43 mm average error against a handheld scanner, while MVS and SfM show 2.32 mm and 7.23 mm [14]; NeRF blurring can merge or split kernels [10]; voxel carving can overestimate volumes due to missed concavities and occlusions [16]. | Representation choice must be validated against ground truth and bias-corrected for trait analysis [10][11][14][16]. |
| Environment and platform motion | Wind below 3.3 m/s leaves frames and point clouds usable, 3.3–5.5 m/s causes noticeable motion noise, and above 5.5 m/s reconstruction becomes unusable [3]; external excitation and plant egomotion degrade scans [13]; mobile platforms provide at least an order of magnitude lower point-cloud quality and resolution than static TLS [13]. | Acquisition scheduling must be environmentally gated, especially in field settings [3][13]. |

The error budget also has a benchmark-transfer caveat [13]. Published shape-completion, denoising, upsampling and reconstruction results are often demonstrated on small-scale clouds with unrealistic properties such as zero-mean Gaussian noise and over-fitted models, so their transfer to real phenotyping data is unproven [13]. A design report should therefore treat benchmark accuracy as a starting hypothesis, not as a guaranteed field error [13].

## Statistical and causal trait analysis

The analysis layer should be multivariate, genotype-aware and uncertainty-aware. In the rice structured-light study, 25 phenotypic traits were extracted from 2000 rice grains across 10 varieties, and filled/unfilled, indica/japonica and variety classification were performed with six machine-learning algorithms: decision tree, random forest, support vector machine, Naive Bayes, XGBoost and BP neural network [1]. The best accuracy for indica and japonica identification was 99.950%, but variety-level identification reached only 47.252% [1]. Thickness had the highest XGBoost feature importance weight of 0.34 for filled/unfilled classification [1]. This pattern is important: high accuracy on coarse biological classes does not imply reliable discrimination among closely related varieties.

Wheat grain analysis adds quality-related traits. A structured-light method extracted 28 3D phenotypic characters and 4 ventral-sulcus traits; length, width, thickness and ventral-sulcus depth MAPEs were 1.83%, 1.86%, 2.19% and 4.81% [6]. A grain-weight model using 32 phenotypic traits from 500 wheat grains achieved cross-validated R² values of 0.77 to 0.83 [6]. WG-3D measured length, width, thickness and volume for 26 Yangmai-series varieties [2]. Maize-ear phenotyping is explicitly positioned for phenotype-to-genotype association analyses, with 1,091 ears of known genotype identity; kernel count reached R² = 0.921 with MAPE = 10.33%, and kernel row number was within two rows for 160 of 168 held-out ears [10]. Kernel row number is described as established early in ear development and exhibiting high broad-sense heritability with consistent QTL associations [10].

Field-scale analysis shows both promise and limits. PhenoRob-F achieved wheat ear precision 0.783, recall 0.822 and mAP 0.853 with YOLOv8m on local field images, lower than its GWHD validation performance of precision 0.924, recall 0.859 and mAP 0.930 [4]. Rice panicle segmentation reached mIoU 0.949 and accuracy 0.987 with SegFormer_B0, maize plant height reached R² = 0.99 against manual measurement, and rapeseed height reached R² = 0.97 [4]. Near-infrared spectral data classified five drought-severity categories with accuracies from 0.977 to 0.996 [4]. Wheat3DGS used 30 views for reconstruction and 6 out-of-distribution views for evaluation, and reported ANOVA F-statistics showing discriminative power for length, width and volume across 42 populations [15]. However, volume estimates were described as either too noisy or capturing different aspects of structure than TLS, and 3DGS was more discriminative than TLS for width and volume in that comparison [15].

The design lesson is that statistical analysis must be coupled to the acquisition controller. If a trait is used for QTL or breeding inference, the system should estimate not only the trait value but also its uncertainty and bias. For example, variety-level classification failure in rice [1], systematic volume underestimation in L1-trained deep models [11], and relative volume proxies in maize [10] all show that trait extraction must include calibration, cross-validation and ground-truth checks.

## Control architecture

A suitable architecture is hierarchical because the cited systems separate platform control, acquisition control and supervisory scheduling [3][4][7][9]. The lowest layer controls physical motion and optical focus, as evidenced by robot speed and camera-height control [4], adaptive visual servoing [8] and autofocus adjustment [5]. The middle layer controls acquisition geometry and reconstruction quality, as evidenced by orthogonal tuning of rotation, scanning angle and stage colour [6] and pose-convergence screening [14]. The top layer supervises scheduling, risk and digital-twin state, as evidenced by DT-FieldPheno’s AHP/fuzzy risk model and heartbeat fault response [3].

The environmental supervisory layer is evidenced by DT-FieldPheno [3]. It uses a five-layer architecture with virtual–physical mapping, model-driven control, multi-protocol coordination and tiered synchronization of heterogeneous data [3]. Its closed loop comprises connection, computation, prediction, decision-making and execution [3]. Virtual collision bodies, trajectory simulation and heartbeat monitoring enable sub-second fault response; track-overrun warning delay was shortened to 1.8 ± 0.2 s, and communication interruption was located in 5 s using a heartbeat mechanism [3]. A dual-layer AHP and fuzzy risk model drives adaptive acquisition scheduling, with wind-speed thresholds of <3.3 m/s for usable data, 3.3–5.5 m/s for noticeable motion noise and >5.5 m/s for unusable reconstruction [3]. Rainfall above 15 mm/h increases lens-splash and connector-leakage risk, with an operational threshold of 14.9 mm/h, and data collection is suspended when PAR is at or below 200 μmol/m²/s [3]. During a 27-day maize field deployment, the system reduced manual inspection workload by 50%, cancelled two high-risk tasks under wind-speed exceedance and optimized two tasks affected by gusts and rainfall [3].

The navigation layer should combine EKF fusion, adaptive visual servoing and predictive control because these methods address pose failure, uncalibrated camera and robot parameters, and unknown wheel–terrain slip [7][8][9]. The loosely coupled EKF provides fault-tolerant pose estimation when GPS or VIO fails [7]. Hybrid visual servoing adapts unknown camera and robot parameters with Lyapunov convergence [8]. NMHE+NMPC handles unknown wheel–terrain interaction and slip, achieving <5 cm tracking error and mean Euclidean error of 0.0423 m after staying on track [9]. PhenoRob-F provides a complementary field-platform reference, with cross-row wheeled mobility, RTK positioning, visual navigation and a maximum speed of 1 m/s [4].

The acquisition layer should use model predictive control over scan speed [2], rotation, scanning angle and stage colour [6], focus [5], illumination and camera height [4], and viewpoint [13][14]. The evidence supplies separate constraints but not a single published joint controller; the separate constraints are scan speed [2], geometry [6], focus [5], image count [11], power [4] and environmental risk [3]. WG-3D quantifies the speed–accuracy trade: 25 mm/s gives length RMSE 0.2256 mm and volume RMSE 1.7740 mm³, while 5 mm/s improves length RMSE to 0.1018 mm and volume RMSE to 0.6457 mm³ but reduces efficiency [2]. Frontiers 2022 quantifies geometric settings: 30° rotation, 37° scanning angle and black stage colour minimize trait error, and excessive rotation reduces efficiency without meaningful accuracy gain [6]. ICCV 2025 quantifies focus convergence: adaptive sharpness mapping reaches >95% in-focus accuracy within five iterations across 35× depth-of-field [5]. ICCV-W 2023 quantifies image count: volume carving needs about 36 images and 18 s per seed, while a deep model reconstructs from 1–3 images with about 2% relative error and negligible inference time [11]. PhenoRob-F quantifies power: halogen illumination consumes more than 1200 W and over 80% of system power, limiting continuous operation on a 100 AH battery to 3–6 h [4].

The separate cost terms can be drawn from the cited measurements: trait uncertainty from placement, registration, mixed-pixel and representation errors [1][11][13][14]; scan time from stage speed and image count [2][11]; energy from illumination power and battery duration [4]; and environmental risk from wind, rain and PAR thresholds [3]. These terms support a proposed MPC cost:

\[
J = \lambda_t U_{\text{trait}} + \lambda_\tau T_{\text{scan}} + \lambda_e E_{\text{energy}} + \lambda_r R_{\text{environment}},
\]

where \(U_{\text{trait}}\) is predicted trait uncertainty [1][11][13][14], \(T_{\text{scan}}\) is acquisition time [2][11], \(E_{\text{energy}}\) is power consumption [4] and \(R_{\text{environment}}\) is environmental risk [3]. The terms are populated from the evidence: \(U_{\text{trait}}\) can be estimated from placement error, registration residual, mixed-pixel score and representation bias [1][11][13][14]; \(T_{\text{scan}}\) is measured by scan speed and image count [2][11]; \(E_{\text{energy}}\) is constrained by illumination power and battery duration [4]; and \(R_{\text{environment}}\) is constrained by wind, rain and PAR thresholds [3].

A reconstruction-quality observer should be added as a design contribution because the sources do not supply a single recursive estimator of point-cloud or trait quality, but they provide ingredients [5][13][14]. Sharpness can be estimated from defocused images [5]. Pose convergence can be screened by ATE [14]. Mixed pixels can be flagged by deviation from the expected Gaussian return distribution [13]. Sampling density can be monitored by neighbour counts in a fixed radius [13]. The observer could therefore output a scalar quality score and trigger re-scanning, viewpoint optimization, filtering or task cancellation [3][5][13][14].

Stability and robustness must be stated carefully [5][7][8][9]. Hybrid visual servoing provides Lyapunov asymptotic convergence for camera and robot parameter adaptation [8]. EKF fusion provides empirical robustness under sensor failure [7]. NMHE+NMPC provides constrained online optimization and field tracking performance [9]. The available evidence does not establish formal frequency-domain robustness margins for a grain-imaging loop, so the design should require empirical validation under texture, illumination, motion and calibration perturbations [5][7][8][9].

## Validation plan and illustrative case studies

Validation should be multi-level: metrological, computational, biological and operational, because the cited studies evaluate accuracy, convergence, biological discrimination and field operation [1][3][4][10][14][15].

| Validation question | Metric | Evidence basis |
| --- | --- | --- |
| Are grain dimensions accurate? | MAPE, RMSE and R² against micrometer or caliper measurements [1][2][6] | Rice structured light used MAPE, RMSE and R² against manual micrometer measurements [1]; WG-3D validated length, width, thickness and volume against electronic vernier calipers [2]; Frontiers wheat used MAPE for length, width, thickness and sulcus depth [6]. |
| Does reconstruction converge? | ATE, average point-cloud error, alignment uncertainty [14][15] | SfM reconstructions were rejected when ATE exceeded 1.5 mm, with 12 of 20 passing [14]; 3DGS and NeRF were compared against a handheld scanner with 0.74 mm and 1.43 mm average accuracy [14]; TLS-to-3DGS alignment was assessed within 10 mm [15]. |
| Are organs detected and segmented? | precision, recall, mAP, IoU, F1 [4][15] | PhenoRob-F reported wheat ear precision 0.783, recall 0.822 and mAP 0.853, and rice panicle mIoU 0.949 [4]; Wheat3DGS reported 3D segmentation IoU and F1 against ground-truth masks [15]. |
| Are traits biologically meaningful? | ANOVA F-statistics, genotype association, heritability, cross-validation R² [6][10][15] | Wheat3DGS reported ANOVA F-statistics for length, width and volume [15]; maize KRN is described as highly heritable with QTL associations [10]; wheat grain weight models achieved R² 0.77–0.83 [6]. |
| Does the controller perform? | tracking error, in-focus accuracy, fault response time [3][5][9] | NMHE+NMPC achieved mean Euclidean error 0.0423 m and <5 cm field tracking [9]; autofocus reached >95% in-focus accuracy within five iterations [5]; DT-FieldPheno achieved sub-second fault response, 1.8 ± 0.2 s track-overrun warning and 5 s communication-fault location [3]. |
| Are environmental limits respected? | wind speed, rainfall, PAR, surface wetness [3][13] | DT-FieldPheno uses wind, rainfall and PAR thresholds [3]; TLS guidance advises avoiding saturated leaves due to rain or dew [13]. |

The case studies show how the architecture should specialize across benchtop, consumer organ-level and field scales [1][2][4][6][10][15].

| Case | Design lesson | Evidence |
| --- | --- | --- |
| Benchtop rice, wheat and corn structured light | Active scanning supports dense point clouds and trait extraction, but placement and variety discrimination remain limiting [1]. | 2000 rice grains across 10 varieties and 100 wheat and corn grains were used; 25 traits were extracted; vertical placement reduced errors; variety accuracy was only 47.252% [1]. |
| WG-3D wheat grain platform | Scan speed is a constrained optimization variable, not a fixed setting [2]. | 25 mm/s gave length RMSE 0.2256 mm and volume RMSE 1.7740 mm³; 5 mm/s improved accuracy but reduced efficiency; back-side data are unavailable [2]. |
| Frontiers wheat grain and sulcus | Acquisition geometry can be tuned by design of experiments [6]. | Orthogonal experiments selected 30° rotation, 37° scanning angle and black stage colour; MAPEs were 1.83%, 1.86%, 2.19% and 4.81% for length, width, thickness and sulcus depth [6]. |
| Maize ear consumer pipeline | Organ-level kernel traits are feasible without specialist hardware, but absolute volume requires caution [10]. | 1,091 ears were processed; kernel count R² was 0.921 and KRN was within two rows for 160/168 ears; volumes are relative proxies not verified against destructive or micro-CT measurement [10]. |
| Wheat field 3DGS | Explicit Gaussian representations can support field-scale head segmentation and trait extraction, but fine structures and volume remain difficult [15]. | Wheat3DGS used 30 views for reconstruction and 6 evaluation views, segmented heads with YOLOv5 and SAM, and reported ANOVA F-statistics, while noting awns were insufficiently reconstructed and volume noisy [15]. |
| PhenoRob-F field robot | Platform control and biological analysis are coupled through resolution, lighting, power and occlusion [4]. | Wheat ear detection, rice panicle segmentation, maize height R² = 0.99 and drought classification were reported, but overlapping ears, nighttime operation and >1200 W lighting were limitations [4]. |

## Limitations and open design questions

The strongest limitation is that 3D phenotyping remains constrained by speed, availability, portability, spatial resolution and cost [16], and the complexity of analysing 3D representations is itself a bottleneck [16]. Active methods can be accurate but costly, while passive methods are cheaper but require heavy computation and can produce lower-quality data [16]. Mobile and UAV platforms can provide more favourable scanning geometry but give at least an order of magnitude lower point-cloud quality and resolution than static TLS [13].

The error budget is fragmented. The literature supplies separate quantified terms—scanner accuracy [1][6], speed-dependent point-cloud error [2], placement error [1], registration rules and failure modes [13], mixed-pixel bias [13], laser-penetration bias and noise [13], reconstruction representation error [11][14] and environmental thresholds [3]—but it does not jointly establish a complete per-trait variance decomposition. The proposed observer and MPC should be treated as a design contribution that assembles these terms, not as a published validated controller.

Trait bias is also unresolved. Volume carving can produce systematic volume error depending on view count [11], deep-learning volume estimates can be systematically underestimated with L1 loss [11], voxel carving can overestimate volumes because of missed concavities and occlusions [16], and maize kernel volumes are described as relative proxies rather than absolute measurements [10]. Similarly, background segmentation can reduce reconstruction accuracy, and NeRF often failed to converge when trained on segmented images [14]. These findings argue for bias-aware validation rather than single-number accuracy reporting.

Platform limitations are material. DT-FieldPheno reports that it has not incorporated soil moisture and canopy microclimate into advanced decision-making, relies on expert-defined AHP weights, lacks robust detection of non-structured anomalies, and does not analyse system-level energy or computational cost; validation is also fully established in maize [3]. PhenoRob-F reports that wheat ear detection struggles with overlapping ears at image edges, that RGB-based yield prediction still needs improvement, that many experiments are conducted at night to avoid changing sunlight, and that halogen illumination dominates power use [4]. These are not minor caveats; they define the operating envelope of the proposed control system.

The most consequential open question is whether a single controller can jointly optimize illumination, exposure, scan trajectory, focus, viewpoint and scheduling against a trait-uncertainty cost. The evidence supports each component separately: scan speed [2], acquisition geometry [6], focus [5], illumination power [4], image count [11], environmental risk [3] and pose estimation [7][9]. What the available evidence does not establish is a published integrated controller that optimizes all of these variables together for grain-trait uncertainty. That is the design gap this report proposes to address.

## Conclusion

The supportable design is a two-tier, closed-loop 3D grain-phenotyping architecture. For kernel-level and sub-millimetre grain traits, the evidence favours active benchtop sensing with controlled stage motion, calibrated illumination, adaptive focus and recursive pose or calibration observers, because structured light and line-laser binocular systems provide the quantified accuracy needed for length, width, thickness and volume [1][6][2]. For head-, ear- and canopy-level traits, passive RGB SfM, MVS, NeRF and 3DGS are viable at breeding scale when combined with segmentation, calibrated poses and environmental scheduling, as shown by maize-ear consumer pipelines, Wheat3DGS field reconstruction and PhenoRob-F field validation [10][15][4]. The modern-control contribution is to connect these layers: EKF and adaptive visual servoing estimate pose and camera parameters [7][8], NMHE+NMPC handles terrain and slip [9], optical sharpness mapping closes focus [5], and a digital twin enforces wind, rain, PAR and fault-response constraints [3]. The decisive uncertainty is not whether the components exist, but whether their separate error terms can be propagated into a validated per-trait uncertainty budget and optimized jointly. Evidence that would change the design judgment would include a published joint MPC over acquisition variables with trait-uncertainty cost, a consolidated variance decomposition from registration and mixed-pixel error into grain-trait error, and long-term field validation beyond the 27-day maize deployment and single-season wheat and maize case studies [3][15][10].

## References

[1] Cereal grain 3D point cloud analysis method for shape extraction and filled/unfilled grain identification based on structured light imaging | Scientific Reports — https://www.nature.com/articles/s41598-022-07221-4
[2] WG-3D: A Low-Cost Platform for High-Throughput Acquisition of 3D Information on Wheat Grain — https://mdpi-res.com/d_attachment/agriculture/agriculture-12-01861/article_deploy/agriculture-12-01861-v4.pdf?version=1669083174
[3] Research on Intelligent Control Technology for a Rail-Based High-Throughput Crop Phenotypic Platform Based on Digital Twins — https://www.mdpi.com/2077-0472/15/11/1217
[4] PhenoRob-F: An autonomous ground-based robot for high-throughput phenotyping of field crops - PMC — https://pmc.ncbi.nlm.nih.gov/articles/PMC12709880/
[5] Optical Model-Driven Sharpness Mapping for Autofocus in Small Depth-of-Field and Severe Defocus Scenarios — https://openaccess.thecvf.com/content/ICCV2025/papers/Fan_Optical_Model-Driven_Sharpness_Mapping_for_Autofocus_in_Small_Depth-of-Field_and_ICCV_2025_paper.pdf
[6] Frontiers | An Intelligent Analysis Method for 3D Wheat Grain and Ventral Sulcus Traits Based on Structured Light Imaging — https://www.frontiersin.org/journals/plant-science/articles/10.3389/fpls.2022.840908/full
[7] Frontiers | A Loosely Coupled Extended Kalman Filter Algorithm for Agricultural Scene-Based Multi-Sensor Fusion — https://www.frontiersin.org/journals/plant-science/articles/10.3389/fpls.2022.849260/full
[8] [2011.01408] Hybrid Visual Servoing Tracking Control of Uncalibrated Robotic Systems for Dynamic Dwarf Culture Orchards Harvest — https://arxiv.org/abs/2011.01408
[9] https://www.roboticsproceedings.org/rss14/p36.pdf — https://www.roboticsproceedings.org/rss14/p36.pdf
[10] Automated Maize Ear Phenotyping Using 3D Reconstructions — https://arxiv.org/pdf/2609.01921.pdf
[11] Deep Learning Based 3d Reconstruction for Phenotyping of Wheat Seeds: a Dataset, Challenge, and Baseline Method — https://openaccess.thecvf.com/content/ICCV2023W/CVPPA/papers/Cherepashkin_Deep_Learning_Based_3d_Reconstruction_for_Phenotyping_of_Wheat_Seeds_ICCVW_2023_paper.pdf
[12] Measuring crops in 3D: using geometry for plant phenotyping | Plant Methods | Springer Nature Link — https://link.springer.com/article/10.1186/s13007-019-0490-0
[13] https://isprs-annals.copernicus.org/articles/X-1-W1-2023/1007/2023/isprs-annals-X-1-W1-2023-1007-2023.pdf — https://isprs-annals.copernicus.org/articles/X-1-W1-2023/1007/2023/isprs-annals-X-1-W1-2023-1007-2023.pdf
[14] High-fidelity wheat plant reconstruction using 3D Gaussian splatting and neural radiance fields - PMC — https://pmc.ncbi.nlm.nih.gov/articles/PMC11945317/
[15] Wheat3DGS: In-field 3D Reconstruction, Instance Segmentation and Phenotyping of Wheat Heads with Gaussian Splatting — https://openaccess.thecvf.com/content/CVPR2025W/V4A/papers/Zhang_Wheat3DGS_In-field_3D_Reconstruction_Instance_Segmentation_and_Phenotyping_of_Wheat_CVPRW_2025_paper.pdf
[16] How to make sense of 3D representations for plant phenotyping: a compendium of processing and analysis techniques - PMC — https://pmc.ncbi.nlm.nih.gov/articles/PMC10288709/


Please begin the extraction now. Output only the JSON list directly, without any chitchat or explanations.