Side-scan sonar remains the workhorse of underwater target detection, yet its images arrive with a familiar compromise: blurred edges where targets end and noise begins. The new LEF-RT-DETR framework takes this problem head-on, not by adding more sensor data, but by reworking how a detection model processes the data it already has. The gains are concrete: a 4.3% improvement in Average Precision and 5.3% in AP50 over the baseline RT-DETR, while cutting parameters by roughly 24% and computational cost by 18%. Those numbers matter because they move real-time detection closer to practical deployment, where every millisecond and every watt of onboard power counts.
The technical choices are worth unpacking because they reflect a clear-eyed reading of the problem. The Gaussian-Edge Enhancement Module directly targets the distorted target contours that plague sonar imagery, while the Multi-Scale Frequency-Spatial Denoising Block separates signal from the acoustic clutter that buries weak returns. The Partial Convolution with Efficient Channel Attention then trims redundant computation without sacrificing the accuracy gains. This is not a flashy redesign; it is incremental, measurable engineering. That is precisely why it works. As we have noted before in Maintaining Reliable Underwater Object Detection Amidst Challenging Conditions, reliability in degraded environments is rarely about one breakthrough, but about stacking careful improvements until the system crosses a usability threshold.
For researchers and operators, the practical implication is straightforward. A lighter, faster model means edge deployment becomes feasible on smaller autonomous underwater vehicles with limited compute budgets. The frequency-domain denoising is particularly promising for cluttered coastal waters, where sediment and turbidity create the kind of noise that defeats conventional spatial filters. This connects to broader efforts in Advancing Underwater Plastic Detection with Integrated Data and AI, where the same challenge of distinguishing targets from background noise limits automated surveys. The common thread is that robust feature extraction, not bigger datasets alone, is what unlocks practical autonomy.
What we would tell a reader asking whether to adopt this approach is simple: watch how the model performs on your own sonar data before committing. The framework was validated on a self-constructed dataset, which is a reasonable starting point but not a guarantee of universal transfer. The real test is whether the edge-frequency fusion holds up across different sonar frequencies, water depths, and bottom types. The open question is whether the parameter reduction sacrifices performance on smaller, lower-contrast targets that are already difficult to resolve. That is the next detail to watch, because real-time detection is only useful when it detects the right things.
