Gas hydrate(GH)is an unconventional resource estimated at 1000-120,000 trillion m^(3)worldwide.Research on GH is ongoing to determine its geological and flow characteristics for commercial produc-tion.After two large-...Gas hydrate(GH)is an unconventional resource estimated at 1000-120,000 trillion m^(3)worldwide.Research on GH is ongoing to determine its geological and flow characteristics for commercial produc-tion.After two large-scale drilling expeditions to study the GH-bearing zone in the Ulleung Basin,the mineral composition of 488 sediment samples was analyzed using X-ray diffraction(XRD).Because the analysis is costly and dependent on experts,a machine learning model was developed to predict the mineral composition using XRD intensity profiles as input data.However,the model’s performance was limited because of improper preprocessing of the intensity profile.Because preprocessing was applied to each feature,the intensity trend was not preserved even though this factor is the most important when analyzing mineral composition.In this study,the profile was preprocessed for each sample using min-max scaling because relative intensity is critical for mineral analysis.For 49 test data among the 488 data,the convolutional neural network(CNN)model improved the average absolute error and coefficient of determination by 41%and 46%,respectively,than those of CNN model with feature-based pre-processing.This study confirms that combining preprocessing for each sample with CNN is the most efficient approach for analyzing XRD data.The developed model can be used for the compositional analysis of sediment samples from the Ulleung Basin and the Korea Plateau.In addition,the overall procedure can be applied to any XRD data of sediments worldwide.展开更多
This study examines the Big Data Collection and Preprocessing course at Anhui Institute of Information Engineering,implementing a hybrid teaching reform using the Bosi Smart Learning Platform.The proposed hybrid model...This study examines the Big Data Collection and Preprocessing course at Anhui Institute of Information Engineering,implementing a hybrid teaching reform using the Bosi Smart Learning Platform.The proposed hybrid model follows a“three-stage”and“two-subject”framework,incorporating a structured design for teaching content and assessment methods before,during,and after class.Practical results indicate that this approach significantly enhances teaching effectiveness and improves students’learning autonomy.展开更多
The big data generated by tunnel boring machines(TBMs)are widely used to reveal complex rock-machine interactions by machine learning(ML)algorithms.Data preprocessing plays a crucial role in improving ML accuracy.For ...The big data generated by tunnel boring machines(TBMs)are widely used to reveal complex rock-machine interactions by machine learning(ML)algorithms.Data preprocessing plays a crucial role in improving ML accuracy.For this,a TBM big data preprocessing method in ML was proposed in the present study.It emphasized the accurate division of TBM tunneling cycle and the optimization method of feature extraction.Based on the data collected from a TBM water conveyance tunnel in China,its effectiveness was demonstrated by application in predicting TBM performance.Firstly,the Score-Kneedle(S-K)method was proposed to divide a TBM tunneling cycle into five phases.Conducted on 500 TBM tunneling cycles,the S-K method accurately divided all five phases in 458 cycles(accuracy of 91.6%),which is superior to the conventional duration division method(accuracy of 74.2%).Additionally,the S-K method accurately divided the stable phase in 493 cycles(accuracy of 98.6%),which is superior to two state-of-the-art division methods,namely the histogram discriminant method(accuracy of 94.6%)and the cumulative sum change point detection method(accuracy of 92.8%).Secondly,features were extracted from the divided phases.Specifically,TBM tunneling resistances were extracted from the free rotating phase and free advancing phase.The resistances were subtracted from the total forces to represent the true rock-fragmentation forces.The secant slope and the mean value were extracted as features of the increasing phase and stable phase,respectively.Finally,an ML model integrating a deep neural network and genetic algorithm(GA-DNN)was established to learn the preprocessed data.The GA-DNN used 6 secant slope features extracted from the increasing phase to predict the mean field penetration index(FPI)and torque penetration index(TPI)in the stable phase,guiding TBM drivers to make better decisions in advance.The results indicate that the proposed TBM big data preprocessing method can improve prediction accuracy significantly(improving R2s of TPI and FPI on the test dataset from 0.7716 to 0.9178 and from 0.7479 to 0.8842,respectively).展开更多
In order to reduce the risk of non-performing loans, losses, and improve the loan approval efficiency, it is necessary to establish an intelligent loan risk and approval prediction system. A hybrid deep learning model...In order to reduce the risk of non-performing loans, losses, and improve the loan approval efficiency, it is necessary to establish an intelligent loan risk and approval prediction system. A hybrid deep learning model with 1DCNN-attention network and the enhanced preprocessing techniques is proposed for loan approval prediction. Our proposed model consists of the enhanced data preprocessing and stacking of multiple hybrid modules. Initially, the enhanced data preprocessing techniques using a combination of methods such as standardization, SMOTE oversampling, feature construction, recursive feature elimination (RFE), information value (IV) and principal component analysis (PCA), which not only eliminates the effects of data jitter and non-equilibrium, but also removes redundant features while improving the representation of features. Subsequently, a hybrid module that combines a 1DCNN with an attention mechanism is proposed to extract local and global spatio-temporal features. Finally, the comprehensive experiments conducted validate that the proposed model surpasses state-of-the-art baseline models across various performance metrics, including accuracy, precision, recall, F1 score, and AUC. Our proposed model helps to automate the loan approval process and provides scientific guidance to financial institutions for loan risk control.展开更多
短期预测在智能电网建设中扮演着重要角色,深刻影响电网发输变配用各个环节的智能化改造。短期预测一般基于系统实测数据,而传感器故障,数据传输错误等原因会导致数据质量下降,严重影响短期预测的精确性。为建立数据质量受损情况下的精...短期预测在智能电网建设中扮演着重要角色,深刻影响电网发输变配用各个环节的智能化改造。短期预测一般基于系统实测数据,而传感器故障,数据传输错误等原因会导致数据质量下降,严重影响短期预测的精确性。为建立数据质量受损情况下的精确短期预测模型,提出了结合数据预处理和双向长短期记忆(bi-directional long short-term memory,Bi-LSTM)的短期预测框架Bi-LSTM-DP(bi-directional long short-term memory data preprocessing)。在Bi-LSTM-DP中,采集的数据首先通过均值填补缺失值,进而基于Savitzky-Golay滤波器对数据降噪,最后采用Bi-LSTM提取时间序列的信息,实现短期预测。为了评估所提方法的性能,文中使用实测的公开数据集分别预测风电发电量和负荷需求,与其他参考方法对比表明了所述方法的有效性和鲁棒性。展开更多
As one of the main methods of microbial community functional diversity measurement, biolog method was favored by many researchers for its simple oper- ation, high sensitivity, strong resolution and rich data. But the ...As one of the main methods of microbial community functional diversity measurement, biolog method was favored by many researchers for its simple oper- ation, high sensitivity, strong resolution and rich data. But the preprocessing meth- ods reported in the literatures were not the same. In order to screen the best pre- processing method, this paper took three typical treatments to explore the effect of different preprocessing methods on soil microbial community functional diversity. The results showed that, method B's overall trend of AWCD values was better than A and C's. Method B's microbial utilization of six carbon sources was higher, and the result was relatively stable. The Simpson index, Shannon richness index and Car- bon source utilization richness index of the two treatments were B〉C〉A, while the Mclntosh index and Shannon evenness were not very stable, but the difference of variance analysis was not significant, and the method B was always with a smallest variance. Method B's principal component analysis was better than A and C's. In a word, the method using 250 r/min shaking for 30 minutes and cultivating at 28 ℃ was the best one, because it was simple, convenient, and with good repeatability.展开更多
基金supported by the Gas Hydrate R&D Organization and the Korea Institute of Geoscience and Mineral Resources(KIGAM)(GP2021-010)supported by the National Research Foundation of Korea(NRF)grant funded by the Korean government(MSIT)(No.2021R1C1C1004460)Korea Institute of Energy Technology Evaluation and Planning(KETEP)grant funded by the Korean government(MOTIE)(20214000000500,Training Program of CCUS for Green Growth).
文摘Gas hydrate(GH)is an unconventional resource estimated at 1000-120,000 trillion m^(3)worldwide.Research on GH is ongoing to determine its geological and flow characteristics for commercial produc-tion.After two large-scale drilling expeditions to study the GH-bearing zone in the Ulleung Basin,the mineral composition of 488 sediment samples was analyzed using X-ray diffraction(XRD).Because the analysis is costly and dependent on experts,a machine learning model was developed to predict the mineral composition using XRD intensity profiles as input data.However,the model’s performance was limited because of improper preprocessing of the intensity profile.Because preprocessing was applied to each feature,the intensity trend was not preserved even though this factor is the most important when analyzing mineral composition.In this study,the profile was preprocessed for each sample using min-max scaling because relative intensity is critical for mineral analysis.For 49 test data among the 488 data,the convolutional neural network(CNN)model improved the average absolute error and coefficient of determination by 41%and 46%,respectively,than those of CNN model with feature-based pre-processing.This study confirms that combining preprocessing for each sample with CNN is the most efficient approach for analyzing XRD data.The developed model can be used for the compositional analysis of sediment samples from the Ulleung Basin and the Korea Plateau.In addition,the overall procedure can be applied to any XRD data of sediments worldwide.
基金2024 Anqing Normal University University-Level Key Project(ZK2024062D)。
文摘This study examines the Big Data Collection and Preprocessing course at Anhui Institute of Information Engineering,implementing a hybrid teaching reform using the Bosi Smart Learning Platform.The proposed hybrid model follows a“three-stage”and“two-subject”framework,incorporating a structured design for teaching content and assessment methods before,during,and after class.Practical results indicate that this approach significantly enhances teaching effectiveness and improves students’learning autonomy.
基金The support provided by the Natural Science Foundation of Hubei Province(Grant No.2021CFA081)the National Natural Science Foundation of China(Grant No.42277160)the fellowship of China Postdoctoral Science Foundation(Grant No.2022TQ0241)is gratefully acknowledged.
文摘The big data generated by tunnel boring machines(TBMs)are widely used to reveal complex rock-machine interactions by machine learning(ML)algorithms.Data preprocessing plays a crucial role in improving ML accuracy.For this,a TBM big data preprocessing method in ML was proposed in the present study.It emphasized the accurate division of TBM tunneling cycle and the optimization method of feature extraction.Based on the data collected from a TBM water conveyance tunnel in China,its effectiveness was demonstrated by application in predicting TBM performance.Firstly,the Score-Kneedle(S-K)method was proposed to divide a TBM tunneling cycle into five phases.Conducted on 500 TBM tunneling cycles,the S-K method accurately divided all five phases in 458 cycles(accuracy of 91.6%),which is superior to the conventional duration division method(accuracy of 74.2%).Additionally,the S-K method accurately divided the stable phase in 493 cycles(accuracy of 98.6%),which is superior to two state-of-the-art division methods,namely the histogram discriminant method(accuracy of 94.6%)and the cumulative sum change point detection method(accuracy of 92.8%).Secondly,features were extracted from the divided phases.Specifically,TBM tunneling resistances were extracted from the free rotating phase and free advancing phase.The resistances were subtracted from the total forces to represent the true rock-fragmentation forces.The secant slope and the mean value were extracted as features of the increasing phase and stable phase,respectively.Finally,an ML model integrating a deep neural network and genetic algorithm(GA-DNN)was established to learn the preprocessed data.The GA-DNN used 6 secant slope features extracted from the increasing phase to predict the mean field penetration index(FPI)and torque penetration index(TPI)in the stable phase,guiding TBM drivers to make better decisions in advance.The results indicate that the proposed TBM big data preprocessing method can improve prediction accuracy significantly(improving R2s of TPI and FPI on the test dataset from 0.7716 to 0.9178 and from 0.7479 to 0.8842,respectively).
文摘In order to reduce the risk of non-performing loans, losses, and improve the loan approval efficiency, it is necessary to establish an intelligent loan risk and approval prediction system. A hybrid deep learning model with 1DCNN-attention network and the enhanced preprocessing techniques is proposed for loan approval prediction. Our proposed model consists of the enhanced data preprocessing and stacking of multiple hybrid modules. Initially, the enhanced data preprocessing techniques using a combination of methods such as standardization, SMOTE oversampling, feature construction, recursive feature elimination (RFE), information value (IV) and principal component analysis (PCA), which not only eliminates the effects of data jitter and non-equilibrium, but also removes redundant features while improving the representation of features. Subsequently, a hybrid module that combines a 1DCNN with an attention mechanism is proposed to extract local and global spatio-temporal features. Finally, the comprehensive experiments conducted validate that the proposed model surpasses state-of-the-art baseline models across various performance metrics, including accuracy, precision, recall, F1 score, and AUC. Our proposed model helps to automate the loan approval process and provides scientific guidance to financial institutions for loan risk control.
文摘短期预测在智能电网建设中扮演着重要角色,深刻影响电网发输变配用各个环节的智能化改造。短期预测一般基于系统实测数据,而传感器故障,数据传输错误等原因会导致数据质量下降,严重影响短期预测的精确性。为建立数据质量受损情况下的精确短期预测模型,提出了结合数据预处理和双向长短期记忆(bi-directional long short-term memory,Bi-LSTM)的短期预测框架Bi-LSTM-DP(bi-directional long short-term memory data preprocessing)。在Bi-LSTM-DP中,采集的数据首先通过均值填补缺失值,进而基于Savitzky-Golay滤波器对数据降噪,最后采用Bi-LSTM提取时间序列的信息,实现短期预测。为了评估所提方法的性能,文中使用实测的公开数据集分别预测风电发电量和负荷需求,与其他参考方法对比表明了所述方法的有效性和鲁棒性。
基金Supported by National and International Scientific and Technological Cooperation Project"The application of Microbial Agents on Mining Reclamation and Ecological Recovery"(2011DFR31230)Key Project of Shanxi academy of Agricultural Science"The Research and Application of Bio-organic Fertilizer on Mining Reclamation and Soil Remediation"(2013zd12)Major Science and Technology Programs of Shanxi Province"Key Technology Research and Demonstration of mining waste land ecosystem Restoration and Reconstruction"(20121101009)~~
文摘As one of the main methods of microbial community functional diversity measurement, biolog method was favored by many researchers for its simple oper- ation, high sensitivity, strong resolution and rich data. But the preprocessing meth- ods reported in the literatures were not the same. In order to screen the best pre- processing method, this paper took three typical treatments to explore the effect of different preprocessing methods on soil microbial community functional diversity. The results showed that, method B's overall trend of AWCD values was better than A and C's. Method B's microbial utilization of six carbon sources was higher, and the result was relatively stable. The Simpson index, Shannon richness index and Car- bon source utilization richness index of the two treatments were B〉C〉A, while the Mclntosh index and Shannon evenness were not very stable, but the difference of variance analysis was not significant, and the method B was always with a smallest variance. Method B's principal component analysis was better than A and C's. In a word, the method using 250 r/min shaking for 30 minutes and cultivating at 28 ℃ was the best one, because it was simple, convenient, and with good repeatability.