You will be provided with a reference and some statements. Please determine whether each statement is 'supported', 'unsupported', or 'unknown' with respect to the reference. Please note:
First, assess whether the reference contains any valid content. If the reference contains no valid information, such as a 'page not found' message, then all statements should be considered 'unknown'.
If the reference is valid, for a given statement: if the facts or data it contains can be found entirely or partially within the reference, it is considered 'supported' (data accepts rounding); if all facts and data in the statement cannot be found in the reference, it is considered 'unsupported'.

You should return the result in a JSON list format, where each item in the list contains the statement's index and the judgment result, for example:
[
    {
        "idx": 1,
        "result": "supported"
    },
    {
        "idx": 2,
        "result": "unsupported"
    }
]

Below are the reference and statements:
<reference>
This ICCV paper is the Open Access version, provided by the Computer Vision Foundation.
Except for this watermark, it is identical to the accepted version;
the final published version of the proceedings is available on IEEE Xplore.

Optical Model-Driven Sharpness Mapping for Autofocus in Small Depth-of-Field
and Severe Defocus Scenarios

1

Chen-Liang Fan1 , Mingpei Cao1,2* , Chih Chien Hung1 , and Yuesheng Zhu2
Foxconn 2 School of Electronic and Computer Engineering, Peking University

r5922507@gmail.com, caomingpei@alumni.pku.edu.cn, chih.chien.hung@ieee.org, zhuys@pku.edu.cn

Abstract

dual-pixel sensors to achieve phase-detection AF [6]. While
effective, these approaches increase system complexity and
cost. Image-based AF techniques have been developed to
overcome these limitations, utilizing high-frequency image
features such as gradients, Laplacians, and variance to estimate sharpness and identify the optimal focus position.
Although adaptable and functional without training under
severe defocus, image-based AF methods often use peaksearching algorithms [40] requiring dense focal sampling.
This is due to the algorithm’s sensitivity to texture and
illumination variations, which causes inconsistent sharpness predictions across different surfaces and lighting conditions, leading to inefficiency in small depth-of-field (DoF)
scenarios where fine-grained adjustments are crucial for accurate focusing.
Recent advances in deep learning (DL) have transformed
autofocus by enabling image-based models that interpret focus cues without extensive sampling. Current DL-based AF
methods typically rely on depth estimation (DE) models in
an end-to-end framework, predicting focus distance from
defocused images and adjusting focus accordingly. Most
are designed for general scenes, like smartphone cameras,
where severe defocus is rare [6, 10]. A few DL-based AF
methods address small DoF applications, such as microscopic bio-imaging [17], but struggle with defocus uncertainty, where the same defocus level occurs at two positions
along the optical axis. This uncertainty limits their working
range, reducing AF reliability under severe defocus.
In this study, we propose a novel autofocus method that
integrates an optical defocus model with a deep learningbased sharpness prediction network. Instead of relying on
direct end-to-end distance estimation, our approach utilizes
optics-based sharpness as an intermediate representation,
leading to more reliable autofocus under severe defocus
conditions. The proposed sharpness indicator is robust to
variations in texture and helps mitigate defocus uncertainty
during focus adjustment, improving accuracy across different surface characteristics. Coupled with an adaptive adjustment algorithm, our method enables efficient autofocus
even from highly defocused starting positions in shallow

Autofocus (AF) is essential for imaging systems, particularly in industrial applications such as automated optical
inspection (AOI), where achieving precise focus is critical. Conventional AF methods rely on peak-searching algorithms that require dense focal sampling, making them
inefficient in small depth-of-field (DoF) scenarios. Deep
learning (DL)-based AF methods, while effective in general
imaging, have a limited working range in small DoF conditions due to defocus uncertainty.
In this work, we propose a novel AF framework that integrates an optical model-based sharpness indicator with
a deep learning approach to predict sharpness from defocused images. We leverage sharpness estimation as a reliable focus measure and apply an adaptive adjustment algorithm to adjust the focus position based on the sharpness-todistance mapping. This method effectively addresses defocus uncertainty and enables robust autofocus across a 35×
DoF range. Experimental results on an AOI system demonstrate that our approach achieves reliable autofocus even
from highly defocused starting points and remains robust
across different textures and illumination conditions. Compared to conventional and existing DL-based approaches,
our method offers improved precision, efficiency, and adaptability, making it suitable for industrial applications and
small DoF scenarios.

1. Introduction
Autofocus (AF) is essential for automatically capturing
high-quality images in applications such as consumer photography, scientific imaging, and automated optical inspection (AOI). Existing AF systems rely on either hardwareassisted or image-based approaches. Hardware-assisted
methods use external sensors such as ultrasonic, LiDAR,
or laser-based range finders to measure object distance and
adjust focus [34, 39], or employ specialized designs like
* Corresponding Author.

6426

DoF environments. The main contributions of this work
can be summarized as follows:
• Propose an optical model-based sharpness indicator capable of representing defocus levels, even in severely defocused images, and utilize a deep learning model to predict
sharpness from defocused inputs.
• Develop an adaptive autofocus adjustment strategy that
leverages the optical model-based sharpness to effectively
overcome defocus uncertainty and reduce the number of
required focus steps across a 35× DoF range.
• Conduct experiments on an AOI system, achieving over
95% in-focus accuracy within five iterations and outperforming baselines in efficiency and adaptability across
various textures and illumination conditions.

this purpose. Both conventional and DL-based DE methods utilize focus and defocus cues, commonly categorized as Depth from Focus (DFF) and Depth from Defocus
(DFD). Conventional DFF methods [5, 8, 21, 23] estimate
depth by identifying the sharpest pixels across focus stacks
but require dense sampling. In contrast, DL-based approaches achieve higher accuracy with fewer sampled focal
planes [9, 20, 32], with Yang et al. further improving performance by integrating Focus Volume [37]. DFD methods
typically use a single defocused image to predict depth. The
U-Net backbone is widely employed for DFD tasks [11],
with Carvalho et al. proposing a U-Net combined with
DenseNet-121 for defocus-based depth estimation [3], a design widely adopted in later studies [19, 22, 38]. Beyond
DFF and DFD, DL-based DE can integrate geometric and
color cues, allowing depth estimation beyond defocus information [1, 13, 18, 41]. However, related studies on DLbased DE for industrial applications remain limited, particularly in addressing severe defocus, small DoF, and environments with minimal geometric and color information.

2. Related Works
2.1. Image based Autofocus
Conventional autofocus (AF) methods assume that objects
approaching the focal plane exhibit sharper high-frequency
features. These methods calculate image sharpness based
on these features, leading to notable variability across different scenes due to dependence on specific image patterns.
Studies have shown that the Tenenbaum Gradient [33]
and Auto-Correlation [29] were among the top-performing
sharpness algorithms. Later methods, such as Tenengrad Variance [24], Laplacian Variance [24], and Modified DCT [16], continued to define sharpness through highfrequency content extraction. Peak-searching techniques
are essential for maximizing sharpness across image stacks
captured at various focus settings. Common approaches include rule-based methods [14], Fibonacci search [15], and
coarse-to-fine search [40], the most robust among conventional AF techniques. However, all searching methods require dense sampling for optimal focus.
Recent deep learning (DL)-based AF methods reduce
sampling by predicting focus distance from defocused images. Herrmann et al. estimated focus distance via classification over 50 intervals [10], while Choi et al. integrated
positional encoding with a classification model for dualpixel imaging [6]. Wang et al. used a CNN to regress distance and image features for AF [35]. DL-based AF has also
been applied to small DoF settings, such as bio-imaging in
microscopy [17, 25, 30, 31], but often suffers from limited
working ranges. Shajkofci et al. demonstrated AF within
an 11× DoF range [31], while Liao et al. achieved one-shot
AF within 7.2× DoF [17]. In contrast, our method achieves
reliable autofocus over a 35× DoF range.

3. Methodology
3.1. Overview
In our method, we first propose a sharpness indicator based
on the defocus optical model, and then utilizing a deep
learning model predicts the sharpness from defocus images.
The sharpness indicator is capable of representing the defocus level of images. Next we propose an adjustment strategy, it utilizes distance mapping results from the sharpness
for position adjusting and finally achieving autofocus, the
framework of the proposed method is presented in Fig. 1.
The defocus model is introduced in Sec. 3.2. The mapping
relation between sharpness and distance is in Sec. 3.3. And
finally the adjustment algorithm is presented in Sec .3.4.

3.2. Optical Model of Defocus
The optical model underlying our AF approach defines
the relationship between sharpness and defocus using the
Circle of Confusion (CoC). In-focus images are produced
when objects lie within the imaging system’s depth-of-field,
where their light rays converge to a point on the sensor
plane. If the object lies outside the DoF, these light rays
form a circular blur, the CoC, on the imaging plane. The
definition of diameter of CoC Dcoc is defined [38]:
  D_{coc} = \frac {|d_{o} - d_{of}|\cdot f}{d_{o}\cdot ( d_{of} - f)\cdot F } \label {"CoC1} 

(1)

The CoC diameter is a function of focal length f , the object distance do , the focal plane of the object dof and finally the aperture F . The relationship with object distance
is sketched in Fig. 2. In particular, the same level of blur
can occur on both sides of the focal plane, introducing uncertainty that presents challenges for both AF and DE tasks.

2.2. Distance Estimation
Distance or depth estimation (DE) predicts depth from single or multiple images and plays a crucial role in DLbased autofocus, despite not being originally designed for

6427

Figure 1. Framework of the proposed autofocus method, integrating an optical model-based sharpness indicator with a deep learning
approach for focus adjustment.

The core of our AF method lies in predicting the relative distance drel between do and dof , where drel =
dof − do . For the convenience, we introduce a constant
f
K = (dof −f
)·F , and refine the Eq. 1 to:
  D_{coc} = \frac {d_{rel} \cdot K}{d_{of} \pm d_{rel}} \label {eq:Dcoc2} 

(2)
Figure 2. Optical model of defocus illustrating CoC variations
across different working distances and the relationship between
DoF and permissible CoC.

The spread of the Circle of Confusion from a point to
a circular blur is the fundamental cause of image defocus.
This effect can be approximated by convolving the in-focus
image with a Gaussian-distributed disk mask, representing
the blur introduced by CoC expansion. The mathematical
representation is given by:
  I_d = I_0 * G(D_{coc}) + noise \label {eq:DefocusedIMG} 

3.3. Sharpness based Distance Estimation
Sharpness Definition: Conventionally, sharpness is defined by high-frequency content within an image, which
is highly sensitive to factors like illumination and textures.
Our method uses the definition of sharpness based on the
CoC model presented in Sec. 3.2. Since CoC directly correlates with the defocus level of images, we define maximum sharpness at the smallest CoC diameter and minimum
sharpness at the largest CoC within the measurement range.
As illustrated in Fig. 3a, the largest CoC occurs on the far
left side of the range. The mathematical definition of sharpness s is given by:

(3)

where Id is the image captured at a position corresponding to drel , and I0 is the image captured at the focal point,
where drel = 0. G(Dcoc ) represents the defocus Gaussian blur kernel, determined by the diameter of the CoC,
Dcoc , and ∗ denoted as convolution. This process acts as
the point-spread function (PSF), a widely used model for
describing defocus phenomena. The DL-model learns to
quantify sharpness from the defocus level.
The range where the CoC is below a certain threshold is
called depth-of-field (DoF). The threshold is the diameter
of the permissible CoC [7], which is defined as:
  CoC_{permissible} = \frac {P_{diagonal}}{N} 

  s = S(d_{rel}) = \frac {{D_{max}} - D_{coc}(d_{rel})}{D_{max} - D_{min} } \label {eq:shaprness} 

(5)

where Dcoc (drel ) is the diameter of the CoC as a function
of drel , which represents the distance to the focal plane. In
the AF applications, the focal plane is assumed to be within
the focus range; therefore, Dmin is 0, corresponding to the
maximum sharpness value of 1, and Dmax corresponds to
the CoC at the boundary on the near side of the lens, where
the sharpness at this boundary is 0. The sharpness value can
represent the defocus level of the images as Fig. 3c shows.
Sharpness Regression: We use a deep learning model to
estimate sharpness as defined in Eq. 5, serving as an intermediate indicator for autofocus. Among tested architectures, U2-Net performed best for sharpness regression. The
detailed experimental comparison is presented in Sec. 4.6.

(4)

where the Pdiagonal represents the diagonal length of the
camera sensor plane, and N is a constant. The technical report [7] provides a data point in which N = 1300. Focus
is achieved when the CoC remains within the permissible
value. Fig. 2 illustrates the permissible CoC for lenses with
identical parameters but varying working distances. The
DoF expands as the working distance increases. Most imaging systems operate at working distances of over a meter,
sometimes even to kilometers, which explains why general
imaging systems rarely encounter severe defocus blur.

6428

For loss function, we empirically determined that a combination of L1 loss and PatchGAN loss [12] provided favorable results for sharpness estimation. This finding aligns
with previous research [4], highlighting that L1 loss and
GAN-based losses each have their merits in related regression tasks. Assume the predicted sharpness ŝi = U (id )
where U is the U2-Net model and id is the defocused image at position drel . The math representation of the loss
function is listed below:
  L_{{L1}} = \sum _{i=1}^{N} \left | s_d - \hat {s}_i \right | 

(6)

  L_{{total}} = L_{{patchGAN}} + constant \cdot L_{{L1}} 

(7)

(a) CoC and Sharpness

(c) Defocus Images and Sharpness

Figure 3. (a) CoC (blue line) and sharpness (red line) derived from
Eq. 5. (b) Range shifted from -5 to 5 mm, where 0 mm means focal
plane. (c) Anodized samples at different distances and sharpness.

Where L_{patchGAN} is the patchGAN loss. sd is the predefine sharpness from Eq. 5 at distance drel . Accurate sharpness estimation is crucial not only for distance estimation
but also as a stop condition for our AF framework.
Distance Mapping: Fig. 3a and Fig. 3b shows the calculated sharpness and an example of estimated sharpness from
the U2-Net model of a specific range. Once the sharpness
of the defocused images is determined, the distance can
be directly mapped through the inverse function of sharpness. The inverse function can be derived from Eq. 2 and
Eq. 5. The Dmin in Eq. 5 is 0. We rewrite Eq. 5 to
Dcoc (drel ) = Dmax · (1 − s). And substitute Eq. 2 to get:
  \frac {d_{rel} \cdot K}{d_{of} \pm d_{rel}} = D_{max} \cdot (1-s)\

Algorithm 1 Autofocus Adjustment Process
Require: Sharpness Model M , Sharpness Inverse Function S −1 , Image img, Max Steps N , Threshold T
1: steps ← 0, prev ← 0, direction ← 1
2: while True do
3:
direction ← 1
4:
sharp ← M (Acquire(img))
5:
dist ← S −1 (sharp, direction)
6:
if steps = 0 then
7:
Adjust focus by dist
8:
else if sharp < prev then
9:
direction ← −1
10:
dist ← S −1 (prev, direction)
11:
Adjust focus by dist
12:
else
13:
Adjust focus by dist
14:
end if
15:
if sharp ≥ T or step >= N then
16:
return
17:
end if
18:
prev ← sharp, steps ← steps + 1
19: end while

(8)

We finally get the inverse function of sharpness s:
  d_{rel} = S^{-1}(s) = \frac {D_{max} \cdot (1-s)\cdot d_{of}}{K \pm D_{max}\cdot (1-s)} \label {eq:inverseS} 

(b) Sharpness and Model Output

(9)

As shown in Sec. 3.2, the same defocus level can occur
on both sides of the focal plane, causing a given sharpness
value to correspond to two possible relative distances drel .
In the next section, we will address direction determination
during focus adjustment.

3.4. Focus Adjustment
Focus adjustment is guided by the distance mapping results
described in Sec. 3.3, with adjustments made based on the
estimated distance. One challenge in this AF process is to
overcome the uncertainty of defocus. To resolve this, we
implement an efficient solution by introducing Take One
More Step: if the sharpness decreases in the second sampling, the adjustment moves away from the focal point; if
it increases, the adjustment moves closer. Although this
approach requires an extra sampling step, it significantly
enhances the adaptability of our model, allowing the DLmodel to focus exclusively on sharpness regression without
the added burden of predicting direction.

In an ideal scenario where distance estimation is perfect,
only a single adjustment would be required. However, due
to inherent model uncertainties and defocus uncertainty in
the optical model, multiple tuning steps are typically needed
until the sharpness reaches a defined threshold. The pseudocode for this process is provided in Alg. 1.
The ± sign in Eq. 9 depends on the moving direction, so
S −1 in Alg. 1 takes sharpness and direction as inputs. Our
adjustment process resolves defocus uncertainty and facilitates convergence within the DoF. This ensures a stable autofocus process across varying defocus conditions.

6429

Table 1. Comparison with Conventional Autofocus

4. Evaluation

Method(Conventional)
Absolute Grad[33]
Variance[33]
Nor-Variance[33]
Auto Correlation[33]
STD Correlation[33]
Entropy[33]
Fisher Entropy[23]
Squared Gradient[33]
Modified DCT[16]
Diagnose Laplacian[5]
Tenengrad Var[24]
Laplacian Var[24]
Brenner Grad[33]
Tenenbaum Grad[33]
Modified Laplacian[33]
Energy Laplacian[33]
Method(Proposed)
Ours
Ours

4.1. Experiment Setup
Experiments were conducted on an Automatic Optical
Inspection (AOI) machine equipped with a COOlLENS
WWK15-110-111 lens and a DaHeng MER2-2000-6GMP camera. This setup provided a depth-of-field (DoF) of
±0.14 mm and a measurement range of ±5 mm, corresponding to 35 times the DoF. Ground truth data was obtained from high-precision position sensors within the motion modules, offering a measurement precision of 1 µm,
ensuring highly accurate reference values and eliminating
the need for synthetic defocus features.
The model was trained and tested on anodized objects,
metal surfaces, printed paper, and printed circuit boards
(PCB), which are common in industrial inspection, covering
a diverse range of material textures to evaluate generalization performance. For both training and testing, focal stacks
were captured across the ±5 mm range with a step size of
0.1 mm, ensuring coverage of varying defocus levels. This
process resulted in 480 focal stacks for training and 800 for
testing, each containing 101 images with the focal plane set
at 0 mm. The images were resized to 512×512 pixels, maintaining sufficient detail while optimizing computational efficiency. The model was trained for 50 epochs using the
Adam optimizer with a learning rate of 2e-5, implemented
on an RTX 4090 GPU.

Size(mm)
0.9
1.2
1.3
0.8
1.3
2.6
0.8
0.8
0.8
0.9
0.9
0.8
0.8
0.9
0.8
0.8
Direction
o
x

o: Initial Direction Correct,

Coarse
6
5
4
7
4
2
7
7
7
6
6
7
7
6
7
7
Coarse
*
*

Fine
4
5
5
3
5
10
3
3
3
4
4
3
3
4
3
3
Fine
*
*

x: Initial Direction Wrong,

Total
10
10
9
10
9
12
10
10
10
10
10
10
10
10
10
10
Total
3
5

In-focus
0.748
0.771
0.892
0.774
0.861
0.855
0.898
0.902
0.93
0.701
0.724
0.874
0.901
0.738
0.858
0.861
In-focus
0.912
0.962

*: Not Applicable

The o and x in Tab. 1 denote the initial adjustment direction of the proposed method. The x means wrong initial
direction and usually requires one additional step to adjust
it, as Fig. 6a shows. The focal plane was manually set to 0
mm by an expert. Given the sampling interval of 0.1 mm
(significantly smaller than the DoF of 0.28 mm) and minor
sample flatness variations, we defined the in-focus region
as (±0.21 mm) (1.5× DoF) around 0 mm. To evaluate AF
performance under severe defocus conditions, we randomly
initialized starting positions within -5 mm to -4 mm or 5
mm to 4 mm across all methods, corresponding to a defocus range of 28x to 35x DoF. All methods were evaluated
on a comprehensive test set to ensure a fair comparison.
Results are presented in Tab. 1. The proposed method
consistently outperformed conventional approaches, even
when the initial adjustment direction was incorrect. Since
only a few conventional methods achieved a success rate
above 90%, we selected the first step at which the success
rate exceeded 90% as the basis for comparing the proposed
and conventional methods. When the initial adjustment direction is correct (o), our approach surpasses 90% in-focus
accuracy by step 3. Even when starting with the wrong initial direction (x), our method outperforms conventional approaches, reaching over 95% in-focus accuracy by step 5.
These results highlight the robustness of our approach in
handling severe defocus conditions, regardless of the initial
adjustment direction.

4.2. Comparison with Conventional Methods
This section compares our autofocus (AF) method with conventional approaches. As mentioned in Sec. 1, conventional
AF methods are sensitive to texture variations. Fig. 4a
shows sharpness curves generated using the NormalizedVariance algorithm (which achieved the best overall score
in a previous study [33]), highlighting this sensitivity. Even
within the same sample domain, peak values and curve
shapes vary significantly, making conventional methods reliant on peak-searching. This variability is the primary
cause of their low efficiency. In contrast, our method computes sharpness using Eq. 5 with a deep learning model,
ensuring greater robustness across different textures. As
shown in Fig. 6a, the peak values and overall curve shapes
remain consistent across sample domains. This stability allows for more effective application of Alg. 1 in focus tuning,
improving efficiency.
In this experiment, we adopt the coarse-to-fine peak
search approach for conventional approaches due to its robustness [40]. In the search method, the coarse step size
should be smaller than the Range, which is defined by the
two local minima surrounding the sharpness peak [33]. To
balance the success rate and sampling time, we fine-tuned
this step size for each algorithm, denoted as size in Tab. 1,
while setting the fine step size to match the DoF.

4.3. Comparison with DL-based Methods
The main limitation of existing DL-based AF methods is
defocus uncertainty. As shown in Fig. 5, defocused images
from the left and right sides of the focal plane appear visually identical, making direct distance estimation prone to
errors. As Fig. 4b illustrates, DE models risk predicting
the distance in the wrong direction, leading to AF failure.

6430

(a) Conventional Sharpness

(b) DE Model Prediction

Figure 4. (a) The sharpness curves from the conventional method
are sensitive to the texture variations. (b) Prediction results from a
DE model. Red is wrong predictions due to defocus uncertainty.

Figure 5. Images of different textures, including anodized objects,
metal, PCB, and printed paper. From left to right: defocused images from the near side to the lens to the far side.

The red data points in Fig. 4b indicate predictions on the
incorrect side of the focal plane. In contrast, the proposed
method determines direction through sharpness change, as
described in Sec. 3.4. This ensures that the risk of incorrect
direction prediction is limited to the initial step.
Initial training of compared models on our dataset resulted in poor performance, with an in-focus rate below
30% in 5 adjustments. To ensure a fair comparison, we
trained and tested the compared models within the same domain for each comparing items (for example, both trained
and tested on the metal dataset), while our method was
trained on the comprehensive dataset (anodized, metal,
printed-paper, PCB), demonstrating better generalization.
We evaluated two models from Depth-from-Defocus studies [22][19] , two from DL-based AF studies [17][10], and
one vision transformer-based depth estimation model [27].
Each method was tested with five adjustment iterations, reporting the highest in-focus rate achieved. Tab. 2 shows that
our method consistently outperforms the compared models
under the small DoF setting at the starting position from severe defocus range.
Table 2. Comparison with DL-based Methods
Methods
Anodized Metal Paper
Nazir et al.[22]
0.846
0.847 0.759
Liao et al.[17]
0.19
0.504 0.628
Lu et al.[19]
0.58
0.33
0.23
Herrmann et al.[10]
0.455
0.83
0.788
Ranftl et al.[27]
0.688
0.305
0.38
Ours(o)
0.998
0.882 0.947
Ours(x)
0.994
0.867 0.912
o: Initial Direction Correct,

Different Textures: The proposed method was evaluated
across different textures. Tab. 3 presents the in-focus rate at
each step, and Fig. 5 shows sample images for each texture. The first row is anodized samples containing rich
high-frequency textures and uniform surface reflectiveness.
The second row is metal samples. Which also have highfrequency texture across the surface. However, its high reflectivity makes it easy to see saturated pixels in images.
The third row is PCB samples, which only have a small
area containing high reflectivity features. The fourth and
fifth rows are printed paper with black content and a white
background. The final row is the average sharpness output
from the U2-Net model. All the samples have different features and are common in industrial inspection tasks.
Table 3. In-focus Rate of Different Textures
Steps
Ano(o)
Ano(x)
Metal(o)
Metal(x)
Paper(o)
Paper(x)
PCB(o)
PCB(x)

PCB
0.859
0.565
0.423
0.835
0.824
0.965
0.919

1
0.598
0.000
0.218
0.000
0.421
0.000
0.663
0.000

2
0.835
0.209
0.841
0.107
0.439
0.430
0.895
0.047

o: Initial Direction Correct,

3
0.996
0.986
0.882
0.867
0.719
0.456
0.965
0.919

4
0.998
0.994
0.945
0.725
0.947
0.623
0.988
0.907

5
1.000
1.000
0.986
0.931
0.833
0.912
0.988
0.977

6
1.000
1.000
0.972
0.960
0.974
0.825
0.977
0.988

7
1.000
0.998
0.972
0.950
0.991
0.904
0.988
0.988

x: Initial Direction Wrong

Our method performs best on anodized and PCB samples due to their rich textures, achieving over 90% in-focus
accuracy by step 3. For metal samples, performance drops
slightly compared to anodized and PCB samples, reaching
85% in-focus accuracy by step 3, likely due to the high reflectivity of metals. Printed paper samples pose the most
significant challenge, achieving 71.9% and 45.6% accuracy
under two initial conditions at step 3. However, both cases
exceed 90% in-focus accuracy in steps 4 and 5. The high
contrast between the white background and black content
may introduce difficulties in sharpness prediction, affecting autofocus performance. Despite these variations, our
method consistently converges to high in-focus accuracy
across different textures.

x: Initial Direction Wrong

4.4. Evaluation of Domain Generalization
This section evaluates the robustness of the proposed autofocus method across different material textures and illumination conditions. The results demonstrate the method’s
ability to maintain high in-focus accuracy across diverse
real-world industrial inspection scenarios from severe defocus starting positions.

6431

Table 4. In-focus Rate of Metal Under Different Illuminations
Steps
Dark(o)
Dark(x)
Normal(o)
Normal(x)
Bright(o)
Bright(x)

1
0.183
0.000
0.366
0.000
0.141
0.000

2
0.866
0.077
0.824
0.127
0.761
0.085

o: Initial Direction Correct,

3
0.923
0.951
0.958
0.944
0.754
0.711

4
0.986
0.761
0.993
0.887
0.852
0.542

5
1
0.993
1
0.986
0.965
0.831

6
0.993
1
0.993
1
0.908
0.93

4.5. AF in All Range and Severe Defocus

7
0.993
1
1
1
0.915
0.859

In this section, we present the detailed results of the proposed AF method on the comprehensive test dataset under
different starting positions. The all range in Tab. 5 corresponds to autofocus starting positions from -5 mm to 5
mm. When the initial direction is correct, the in-focus rate
reaches 64.3% at the first step, 90% at step 3, and exceeds
95% at step 5. When the initial direction is incorrect, performance is significantly worse in the first two steps but eventually reaches 95% at step 5.
For the severe defocus range (28× to 35× DoF), corresponding to initial positions between -5 to -4 mm and 4 to
5 mm, our method achieves 90% and 85% in-focus rates
at step 3, respectively, and eventually exceeds 95%. This
performance closely matches that of the all range case,
demonstrating that our method remains equally effective
even from highly defocused starting points. Overall, the
proposed method maintains high autofocus accuracy even
under severe defocus conditions, with no significant drop
in performance compared to the all-range case. Achieving
95% in-focus accuracy at step 5 in both cases demonstrates
its suitability for small DoF scenarios such as industrial inspection and microscopy.

x: Initial Direction Wrong

Different Illuminations: To evaluate the robustness of
our method under varying illumination conditions, we conducted experiments on metal samples due to their high reflectivity. We collected three sets of samples under bright
(average intensity 168), normal (128), and dark (88) conditions. The predicted sharpness values for each condition are
shown in Fig. 6b. Under bright conditions, some predictions
appear abnormal, likely due to image saturation.
Since our method derives the adjustment distance from
the inverse function of sharpness (Eq. 9), variations in
sharpness predictions directly impact autofocus performance. As shown in Tab. 4, the in-focus rate under bright
conditions is the lowest, reaching only 70% by step 3,
whereas both normal and dark conditions exceed 90% at the
same step. We assume that under bright illumination, saturation effects disrupt sharpness prediction, leading to occasional AF failures, as shown in Fig. 6b. However, by steps
5 and 6, the in-focus rate still reaches 90% across all conditions, demonstrating the robustness of our method against
illumination variations and reflective surfaces. This finding
demonstrates the method’s capability to address common
challenges in industrial environments.

Table 5. In-focus Rate of All and Severe Defocus (SD) Range
Steps
All(o)
All(x)
SD(o)
SD(x)

1
0.643
0.027
0.428
0.000

2
0.807
0.521
0.781
0.169

o: Initial Direction Correct,

3
0.939
0.860
0.912
0.872

4
0.978
0.896
0.969
0.843

5
0.967
0.970
0.975
0.962

6
0.981
0.949
0.976
0.965

7
0.985
0.964
0.982
0.961

x: Initial Direction Wrong,

4.6. Evaluation of Sharpness Estimation
Optical model-based sharpness estimation is crucial for autofocus, as focus adjustment relies on the inverse function
of sharpness (Eq. 9). This section examines the effectiveness of deep learning-based sharpness regression, comparing different model architectures and analyzing how texture
and illumination conditions influence sharpness prediction
and autofocus performance.
We evaluate several deep learning architectures for
sharpness regression, including VIT [27], U-Net [28], ResUnet [36], Dense-Unet, and U2-Net [26]. The CNN-based
models follow a five-layer encoder-decoder structure, while
VIT utilizes a self-attention mechanism. Experiments on
800 focal stacks show that U2-Net achieved the best overall
performance, as Tab. 6 shown. The details of the evaluation
metrics can be found in Caden et al. [2].
Fig. 6a illustrates the relation between predicted sharpness and the distance across different textures. While deviations from the ground truth increase with severe defocus, sharpness values consistently converge near the focal
plane regardless of texture. This consistency confirms that

(a) Sharpness Curves of Different Textures and the AF Process

(b) Different Illumination and the AF Process

Figure 6. (a) Consistent sharpness across textures results in a stable AF process. (b) Illumination variations in metal samples, particularly under bright conditions, lead to slightly less stable sharpness and AF performance.

6432

U-Net
Res-Unet
Dense-Unet
U2-Net
VIT

MAE
0.0582
0.06
0.0607
0.057
0.0639

Table 6. Sharpness Prediction Performance Across Different Models
MSE
RMSE LogMSE LogRMSE AbsRel SqrRel
δ1 ↑
0.008 0.0585
0.0009
0.0188
0.0429 0.0061 0.732
0.0081 0.0606 0.00089
0.0194
0.0435 0.0059 0.708
0.0085 0.0608 0.00095
0.0194
0.0443 0.0063 0.7224
0.0076 0.0573 0.00083
0.0182
0.0416 0.0057 0.7322
0.0089 0.0657
0.001
0.0208
0.0486
0.007 0.7194

δ2 ↑
0.8431
0.8283
0.8356
0.8508
0.8451

δ3 ↑
0.8927
0.8865
0.8878
0.9028
0.904

Table 7. Sharpness Evaluation Across Different Textures and Illumination Using U2-Net

Ano
Metal
Paper
PCB

MAE
0.0374
0.0601
0.1166
0.0348

MSE
0.0026
0.0083
0.0227
0.0033

Bright
Normal
Dark

0.0839
0.0418
0.0540

0.0147
0.0038
0.0062

Different Textures
RMSE LogMSE LogRMSE AbsRel SqrRel
0.0375
0.0003
0.0119
0.0269 0.0019
0.0603
0.0009
0.0191
0.0456 0.0065
0.1169
0.0025
0.0376
0.0820 0.0162
0.0359
0.0003
0.0112
0.0254 0.0025
Different Illumination of Metal Samples
0.0843
0.0015
0.0263
0.0629 0.0112
0.0420
0.0004
0.0135
0.0317 0.0030
0.0541
0.0007
0.0174
0.0418 0.0052

δ1 ↑
0.8052
0.7250
0.4878
0.8399

δ2 ↑
0.9106
0.8499
0.6476
0.9186

δ3 ↑
0.9511
0.9053
0.7348
0.9497

0.6478
0.7874
0.7420

0.7966
0.8918
0.8629

0.8698
0.9350
0.9121

5. Conclusion

the sharpness trends remain stable across different surfaces,
making depth estimation more predictable. However, inaccurate sharpness predictions lead to incorrect distance estimations, affecting autofocus performance.

In this work, we proposed a novel autofocus (AF) method
to address the challenges associated with small DoF scenarios and severe defocus conditions. By integrating an optical model-based sharpness indicator with a deep learning
framework, our approach enables distance estimation even
from severely defocused images through sharpness-distance
mapping, achieving autofocus without dense sampling and
avoiding defocus uncertainty across a 35× DoF range.

Texture-dependent sharpness estimation performance,
detailed in Tab. 7, directly impacts autofocus success rates.
Anodized and PCB samples exhibit the lowest estimation errors with higher in-focus rates and fewer adjustment steps. In contrast, metal samples show moderately
higher errors due to reflective surfaces, while printed paper samples present the most significant challenges as highcontrast regions disrupt sharpness consistency. As shown in
Tab. 3, these less reliable estimates require additional refinements to reach acceptable focus levels; overall, however, the
proposed approach achieves robust autofocus performance
across all textures within a few steps.

Extensive evaluations demonstrate that our method outperforms conventional image-based and existing DL-based
AF techniques in accuracy, efficiency, and adaptability.
On a comprehensive test dataset containing anodized objects, metal surfaces, printed paper, and PCB, our method
achieves over 95% in-focus accuracy within five iterations, effectively handling highly defocused starting positions while maintaining stability across diverse material surfaces common in industrial inspection. This highlights its
suitability for small DoF scenarios such as industrial inspection and microscopy. By overcoming key limitations of existing AF methods, our approach enhances autofocus reliability in applications requiring precise focus control. Future
work will explore its extension to dynamic environments
and multi-plane focusing systems to further improve its capabilities.

Illumination conditions, especially in highly reflective
samples, also affect sharpness estimation. In bright settings,
saturation effects degrade model performance. As shown
in Tab. 4 in the illumination section, metal samples under
bright lighting exhibit the highest prediction errors, aligning
with our assumption in Sec. 4.4 that overexposure reduces
sharpness reliability.
Overall, the accuracy of sharpness estimation is directly
linked to autofocus performance. The results confirm that
U2-Net provides the most reliable sharpness predictions
across different textures and lighting conditions, improving autofocus efficiency. Poor sharpness regression, particularly in highly reflective or high-contrast surfaces, negatively impacts autofocus performance by increasing the required adjustment steps. These findings emphasize the importance of robust sharpness estimation in ensuring precise
and efficient autofocus in small DoF imaging systems.

Acknowledgements
This research was funded by the MCEG Business Group,
Foxconn. All intellectual property rights arising from this
study are owned by Foxconn. We also thank Hao-Cheng
Kan and Chun-Jen Weng for valuable discussions.

6433

References

[12] Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A.
Efros. Image-to-image translation with conditional adversarial networks. In 2017 IEEE Conference on Computer Vision
and Pattern Recognition, CVPR 2017, Honolulu, HI, USA,
July 21-26, 2017, pages 5967–5976. IEEE Computer Society, 2017. 4
[13] Hyungjoo Jung, Youngjung Kim, Dongbo Min, Changjae
Oh, and Kwanghoon Sohn. Depth prediction from a single
image with conditional adversarial networks. In 2017 IEEE
International Conference on Image Processing (ICIP), pages
1717–1721. IEEE, 2017. 2
[14] Nasser Kehtarnavaz and H-J Oh. Development and real-time
implementation of a rule-based auto-focus algorithm. RealTime Imaging, 9(3):197–203, 2003. 2
[15] Eric Krotkov. Focusing. International Journal of Computer
Vision, 1(3):223–237, 1988. 2
[16] Sang-Yong Lee, Yogendera Kumar, Ji-Man Cho, Sang-Won
Lee, and Soo-Won Kim. Enhanced autofocus algorithm using robust focus measure and fuzzy reasoning. IEEE Transactions on Circuits and Systems for Video Technology, 18(9):
1237–1246, 2008. 2, 5
[17] Jun Liao, Xu Chen, Ge Ding, Pei Dong, Hu Ye, Han
Wang, Yongbing Zhang, and Jianhua Yao. Deep learningbased single-shot autofocus method for digital microscopy.
Biomedical Optics Express, 13(1):314–327, 2021. 1, 2, 6
[18] Ce Liu, Suryansh Kumar, Shuhang Gu, Radu Timofte, and
Luc Van Gool. Single image depth prediction made better:
A multivariate gaussian take. In IEEE/CVF Conference on
Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 17346–17356.
IEEE, 2023. 2
[19] Yawen Lu, Garrett Milliron, John Slagter, and Guoyu Lu.
Self-supervised single-image depth estimation from focus
and defocus clues. IEEE Robotics and Automation Letters, 6
(4):6281–6288, 2021. 2, 6
[20] Maxim Maximov, Kevin Galim, and Laura Leal-Taixé. Focus on defocus: Bridging the synthetic to real domain gap for
depth estimation. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle,
WA, USA, June 13-19, 2020, pages 1068–1077. Computer
Vision Foundation / IEEE, 2020. 2
[21] Shree K Nayar and Yasuo Nakagawa. Shape from focus.
IEEE Transactions on Pattern analysis and machine intelligence, 16(8):824–831, 1994. 2
[22] Saqib Nazir, Lorenzo Vaquero, Manuel Mucientes, Vı́ctor M
Brea, and Daniela Coltuc. Depth estimation and image
restoration by deep learning from defocused images. IEEE
Transactions on Computational Imaging, 9:607–619, 2023.
2, 6
[23] Pavel Pavliček and Ivana Hamarová. Shape from focus for
large image fields. Applied Optics, 54(33):9747–9751, 2015.
2, 5
[24] José Luis Pech-Pacheco, Gabriel Cristóbal, Jesús ChamorroMartinez, and Joaquı́n Fernández-Valdivia. Diatom autofocusing in brightfield microscopy: a comparative study.
In Proceedings 15th International Conference on Pattern
Recognition. ICPR-2000, pages 314–317. IEEE, 2000. 2,
5

[1] Alexey Bochkovskiy, Amaël Delaunoy, Hugo Germain,
Marcel Santos, Yichao Zhou, Stephan R. Richter, and
Vladlen Koltun. Depth pro: Sharp monocular metric depth in
less than a second. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore,
April 24-28, 2025. OpenReview.net, 2025. 2
[2] Cesar Cadena, Yasir Latif, and Ian D Reid. Measuring
the performance of single image depth estimation methods.
In 2016 IEEE/RSJ International Conference on Intelligent
Robots and Systems (IROS), pages 4150–4157. IEEE, 2016.
7
[3] Marcela Carvalho, Bertrand Le Saux, Pauline TrouvéPeloux, Andrés Almansa, and Frédéric Champagnat. Deep
depth from defocus: how can defocus blur improve 3d estimation using dense neural networks? In Proceedings of the
European Conference on Computer Vision (ECCV) Workshops, 2018. 2
[4] Marcela Carvalho, Bertrand Le Saux, Pauline TrouvéPeloux, Andrés Almansa, and Frédéric Champagnat. On regression losses for deep depth estimation. In 2018 25th IEEE
International Conference on Image Processing (ICIP), pages
2915–2919. IEEE, 2018. 4
[5] Chih-Yen Chen, Lijuan Wang, Chen-Liang Fan, Pi-Ying
Cheng, and Chun-Jen Weng. Surface profiling of object using varifocal lens with image contrast. Optical Review, 26
(5):493–499, 2019. 2, 5
[6] Myungsub Choi, Hana Lee, and Hyong-Euk Lee. Exploring
positional characteristics of dual-pixel data for camera autofocus. In IEEE/CVF International Conference on Computer
Vision, ICCV 2023, Paris, France, October 1-6, 2023, pages
13112–13122. IEEE, 2023. 1, 2
[7] Toshiba Teli Corporation. Consideration for depth of field in
machine vision, 2016. White Paper, DAA00940A. Copyright
©2016-2020 Toshiba Teli Corporation, All rights reserved. 3
[8] Chen-Liang Fan, Chun-Jen Weng, Yu-Hsin Lin, and PiYing Cheng. Surface profiling measurement using varifocal
lens based on focus stacking. In 2018 IEEE International
Instrumentation and Measurement Technology Conference
(I2MTC), pages 1–5. IEEE, 2018. 2
[9] Caner Hazirbas, Sebastian Georg Soyer, Maximilian Christian Staab, Laura Leal-Taixé, and Daniel Cremers. Deep
depth from focus. In Computer Vision–ACCV 2018: 14th
Asian Conference on Computer Vision, Perth, Australia, December 2–6, 2018, 2018. 2
[10] Charles Herrmann, Richard Strong Bowen, Neal Wadhwa,
Rahul Garg, Qiurui He, Jonathan T. Barron, and Ramin
Zabih. Learning to autofocus. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR
2020, Seattle, WA, USA, June 13-19, 2020, pages 2227–
2236. Computer Vision Foundation / IEEE, 2020. 1, 2, 6
[11] Hayato Ikoma, Cindy M Nguyen, Christopher A Metzler, Yifan Peng, and Gordon Wetzstein. Depth from defocus with
learned optics for imaging and occlusion-aware depth estimation. In 2021 IEEE International Conference on Computational Photography (ICCP), pages 1–12. IEEE, 2021. 2

6434

[25] Henry Pinkard, Zachary Phillips, Arman Babakhani,
Daniel A Fletcher, and Laura Waller. Deep learning for
single-shot autofocus microscopy. Optica, 6(6):794–797,
2019. 2
[26] Xuebin Qin, Zichen Zhang, Chenyang Huang, Masood Dehghan, Osmar R Zaiane, and Martin Jagersand. U2-net: Going deeper with nested u-structure for salient object detection. Pattern recognition, 106:107404, 2020. 7
[27] René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In 2021 IEEE/CVF
International Conference on Computer Vision, ICCV 2021,
Montreal, QC, Canada, October 10-17, 2021, pages 12159–
12168. IEEE, 2021. 6, 7
[28] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. Unet: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted
intervention–MICCAI 2015: 18th international conference,
Munich, Germany, October 5-9, 2015, proceedings, part III
18, pages 234–241. Springer, 2015. 7
[29] Andrés Santos, Carlos ORTIZ DE SOLÓRZANO, Juan José
Vaquero, Javier Márquez Pena, Norberto Malpica, and Francisco del Pozo. Evaluation of autofocus functions in molecular cytogenetic analysis. Journal of microscopy, 188(3):264–
272, 1997. 2
[30] Philipp Johannes Schubert, Rangoli Saxena, and Joergen Kornfeld. Deepfocus: Fast focus and astigmatism correction for
electron microscopy. Nature Communications, 15(1):948,
2024. 2
[31] Adrian Shajkofci and Michael Liebling. Deepfocus: a fewshot microscope slide auto-focus using a sample invariant
cnn-based sharpness function. In 2020 IEEE 17th International Symposium on Biomedical Imaging (ISBI), pages 164–
168. IEEE, 2020. 2
[32] Haozhe Si, Bin Zhao, Dong Wang, Yunpeng Gao, Mulin
Chen, Zhigang Wang, and Xuelong Li. Fully self-supervised
depth estimation from defocus clue. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR
2023, Vancouver, BC, Canada, June 17-24, 2023, pages
9140–9149. IEEE, 2023. 2
[33] Yu Sun, Stefan Duthaler, and Bradley J Nelson. Autofocusing in computer microscopy: selecting the optimal focus
algorithm. Microscopy research and technique, 65(3):139–
149, 2004. 2, 5
[34] Frode Tjontveit, Erik Hellerud, Kristian Tangeland, and Oystein Damhaug. Lidar-assisted auto focus to enable snapshots
of close objects. Technical Disclosure Commons, 2020. 1
[35] Chengyu Wang, Qian Huang, Ming Cheng, Zhan Ma, and
David J Brady. Deep learning for camera autofocus. IEEE
Transactions on Computational Imaging, 7:258–271, 2021.
2
[36] Xiao Xiao, Shen Lian, Zhiming Luo, and Shaozi Li.
Weighted res-unet for high-quality retina vessel segmentation. In 2018 9th international conference on information
technology in medicine and education (ITME), pages 327–
331. IEEE, 2018. 7
[37] Fengting Yang, Xiaolei Huang, and Zihan Zhou. Deep depth
from focus with differential focus volume. In IEEE/CVF

Conference on Computer Vision and Pattern Recognition,
CVPR 2022, New Orleans, LA, USA, June 18-24, 2022,
pages 12632–12641. IEEE, 2022. 2
[38] Anmei Zhang and Jian Sun. Joint depth and defocus estimation from a single image using physical consistency. IEEE
Transactions on Image Processing, 30:3419–3433, 2021. 2
[39] Xiaobo Zhang, Fumin Fan, Mehdi Gheisari, and Gautam Srivastava. A novel auto-focus method for image processing using laser triangulation. Ieee Access, 7:64837–64843, 2019.
1
[40] Yupeng Zhang, Liyan Liu, Weitao Gong, Haihua Yu, Wei
Wang, Chongying Zhao, Peng Wang, and Toshitsugu Ueda.
Autofocus system and evaluation methodologies: A literature review. Sensors & Materials, 30, 2018. 1, 2, 5
[41] Zixiang Zhao, Jiangshe Zhang, Xiang Gu, Chengli Tan,
Shuang Xu, Yulun Zhang, Radu Timofte, and Luc Van Gool.
Spherical space feature decomposition for guided depth map
super-resolution. In IEEE/CVF International Conference on
Computer Vision, ICCV 2023, Paris, France, October 1-6,
2023, pages 12513–12524. IEEE, 2023. 2

6435
</reference>

<statements>
1. The published grain, field and reconstruction pipelines justify treating the scanner, motion stage or robot, illumination, focus, reconstruction software and trait-analysis layer as one controllable dynamic system.
2. The evidence supports three load-bearing conclusions, each developed below from quantified pipeline results and control results.
3. optical-model sharpness mapping closes autofocus to more than 95% in-focus accuracy within five iterations over a 35× depth-of-field range
4. The design contribution is therefore not to invent a new sensor, but to couple these methods into a supervisory digital twin that optimizes scan speed, viewpoint, focus, illumination and scheduling against a trait-uncertainty cost, while explicitly budgeting registration, mixed-pixel, laser-penetration, occlusion and benchmark-transfer errors that the literature quantifies only separately
5. The measured dependencies in the cited pipelines—stage speed, acquisition geometry, focus, illumination and camera height, and plant motion—support abstracting the controllable system as
6. \(u\) is the manipulated acquisition vector
7. focus is a dynamic optical state
8. an optical defocus model gives a sharpness-to-distance map for autofocus
9. recursive platform and optical state estimation
10. For the optical channel, autofocus can be modelled as a state-estimation problem
11. The defocused image is represented as \(I_d=I_0*G(D_{\text{coc}})+\text{noise}\), where a Gaussian blur kernel encodes defocus
12. A deep network predicts sharpness from defocused images, and an adaptive adjustment algorithm uses a sharpness-to-distance map to move the focus stage
13. The method explicitly addresses defocus uncertainty, where the same defocus level occurs at two axial positions, and achieves over 95% in-focus accuracy within five iterations across a 35× depth-of-field range
14. Performance varies by texture and illumination: printed paper is harder than anodized or PCB samples, and bright conditions can cause saturation-related prediction errors
15. The lowest layer controls physical motion and optical focus, as evidenced by autofocus adjustment
16. The acquisition layer should use model predictive control over focus
17. The evidence supplies separate constraints but not a single published joint controller; the separate constraints are focus
18. ICCV 2025 quantifies focus convergence: adaptive sharpness mapping reaches >95% in-focus accuracy within five iterations across 35× depth-of-field.
19. A reconstruction-quality observer should be added as a design contribution because the sources do not supply a single recursive estimator of point-cloud or trait quality, but they provide ingredients
20. Sharpness can be estimated from defocused images
21. The observer could therefore output a scalar quality score and trigger re-scanning, viewpoint optimization, filtering or task cancellation
22. Stability and robustness must be stated carefully
23. The available evidence does not establish formal frequency-domain robustness margins for a grain-imaging loop, so the design should require empirical validation under texture, illumination, motion and calibration perturbations
24. Does the controller perform? tracking error, in-focus accuracy, fault response time
25. autofocus reached >95% in-focus accuracy within five iterations
26. The evidence supports each component separately: scan speed [2], acquisition geometry [6], focus [5]
27. The modern-control contribution is to connect these layers: EKF and adaptive visual servoing estimate pose and camera parameters [7][8], NMHE+NMPC handles terrain and slip [9], optical sharpness mapping closes focus [5]
</statements>

Begin the assessment now. Output only the JSON list, without any conversational text or explanations.