高级检索+

基于声振时频图像和MobileNetV3-RA的脆柿硬度预测方法

Research on firmness prediction method of crisp persimmons based on acoustic vibration time-frequency images and MobileNetV3-RA

  • 摘要: 为解决传统方法难以实现脆柿硬度无损预测,以及现有声振检测方法对非平稳响应信号特征利用不足的问题,该研究将时频信号二维表征方法与MobileNetV3-RA模型结合用于脆柿硬度的预测。首先,利用自制的声学振动检测装置采集脆柿声振响应信号,通过马尔可夫转移场(Markov transition field, MTF)、格拉姆角差分场(Gramian angular difference field,GADF)、连续小波变换(continuous wavelet transform, CWT)和离散S变换(discrete s-transform, DST),将一维声振响应信号转换为256×256像素的二维时频图像。随后,以二维时频图像作为输入构建了包含挤压-激发(squeeze-and-excitation, SE)模块并和残差注意力(residual attention, RA)模块的MobileNetV3-RA模型用于脆柿硬度预测,并将其与常规的二维卷积神经网络(two-dimensional convolutional neural network,2D-CNN)和18层残差网络(18-layer Residual network,ResNet18)模型进行对比。结果表明,以CWT图像作为输入的MobileNetV3-RA获得了最佳的预测效果,预测集上的决定系数RP2、均方根误差RMSEP和相对分析误差RPDp分别为0.909、0.885 N/mm和3.307,且其在模型参数量、单次推理时间等方面也显著优于传统的2D-CNN和ResNet18模型。综上,CWT时频二维表征与MobileNetV3-RA相结合能够在提高脆柿硬度预测精度的同时显著降低模型的计算和存储资源需求,可为水果硬度的快速无损检测提供方法新方法与新思路。

     

    Abstract: Conventional penetration testing is destructive and therefore cannot meet the demand for rapid and nondestructive prediction of crisp persimmon firmness. Moreover, existing acoustic-vibration methods do not fully exploit the characteristics of nonstationary response signals. To address these limitations, this study developed a firmness prediction method that combines two-dimensional representations of acoustic-vibration signals with a lightweight MobileNetV3-RA model. First, 570 crisp persimmons were collected from Gongcheng County, Guangxi Zhuang Autonomous Region, China, covering four harvest batches. A sinusoidal swept-frequency excitation signal ranging from 100 to 1,500 Hz was applied to each fruit, and its acoustic-vibration response was recorded using a microphone at a sampling frequency of 8,192 Hz for 5 s. Reference firmness was measured by a texture analyzer using a penetration test with a cylindrical probe 6 mm in diameter, a penetration speed of 1 mm/s, and a penetration depth of 8 mm. Second, four methods, namely the Markov transition field (MTF), Gramian angular difference field (GADF), continuous wavelet transform (CWT), and discrete S-transform (DST), were used to convert each one-dimensional vibration response into a 256 × 256-pixel two-dimensional image. These four representations characterize the response from complementary perspectives, including temporal state transitions, angular correlations, multiscale variations, and frequency-dependent time localization, thereby allowing the suitability of each representation for firmness-related feature learning to be examined. A MobileNetV3-RA regression model was then constructed by retaining the squeeze-and-excitation (SE) mechanism of MobileNetV3 and introducing a residual attention (RA) module. The lightweight backbone was used to reduce model complexity, whereas the SE and RA modules were used to strengthen the extraction of channel and spatial features from the two-dimensional images. The samples were divided into training, validation, and prediction sets at a ratio of 6:2:2, corresponding to 342, 114, and 114 samples, respectively. A conventional two-dimensional convolutional neural network (2D-CNN) and an 18-layer residual network (ResNet18) were selected as comparison models. All models were evaluated using the same sample allocation. Model performance was assessed using the coefficient of determination for the prediction set ( R_\textp^2 ), root mean square error of prediction (RMSEP), and residual predictive deviation (RPDP). The results showed that CWT and DST generally provided better predictive performance than GADF and MTF, and the combination of CWT images and MobileNetV3-RA achieved the best overall result among the tested model–representation combinations. With CWT images as input, MobileNetV3-RA obtained an R_\textp^2 of 0.909, an RMSEP of 0.885 N/mm, and an RPDP of 3.307 on the prediction set. Under the same CWT input condition, the corresponding values for 2D-CNN were 0.843, 1.159 N/mm, and 2.526, whereas those for ResNet18 were 0.804, 1.258 N/mm, and 2.277. Compared with 2D-CNN, MobileNetV3-RA increased R_\textp^2 by 0.066 and reduced RMSEP by 0.274 N/mm; compared with ResNet18, it increased R_\textp^2 by 0.105 and reduced RMSEP by 0.373 N/mm. The parameter count, model storage size, floating-point operations, and mean single-sample inference time of MobileNetV3-RA were 0.80 million, 3.00 MB, 0.14 G, and 8.5 ms, respectively. These four indicators were reduced by 90.6%, 90.8%, 41.7%, and 30.3%, respectively, relative to 2D-CNN, and by 92.9%, 93.3%, 92.2%, and 45.5%, respectively, relative to ResNet18. In summary, the proposed combination of CWT-based two-dimensional acoustic-vibration representation and MobileNetV3-RA improved crisp persimmon firmness prediction while substantially reducing computational and storage requirements under the present experimental conditions. The method provides a technical reference for rapid and nondestructive fruit firmness assessment and has potential for subsequent validation in embedded devices and online sorting systems.

     

/

返回文章
返回