Adaptive hyperparameter optimization for author name disambiguation.
In the process of author name disambiguation (AND), varying characteristics and noise of different blocks significantly impact disambiguation performance. In this paper, we propose a block‐based adaptive hyperparameter optimization method that assigns optimal hyperparameters to each block without al...
| Publicado en: | Journal of the Association for Information Science & Technology Vol. 76; no. 8; pp. 1082 - 1105 |
|---|---|
| Autores principales: | , |
| Formato: | equations & formulas research tables/charts Journal Article |
| Publicado: |
Wiley-Blackwell
Aug2025
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=187574201&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 187574201 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 23301635 H6JN jtl: Journal of the Association for Information Science & Technology issn: 23301635 maglogo: N pubinfo: dt: Aug2025 vid: 76 iid: 8 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 187574201 183788831 187574201 187574201 10.1002/asi.24996 187574201 ppf: 1082 ppct: 23 formats: tig: atl: Adaptive hyperparameter optimization for author name disambiguation. aug: au: Lu, Shuo Zhou, Yong affil: School of Management, Xi'an University of Architecture and Technology, Xi'an City Shaanxi Province, , China sug: subj: Information Storage Information Retrieval Authorship Models, Statistical Algorithms Human Data Analysis, Computer Assisted Machine Learning Random Forest Data Mining Regression Cluster Analysis Probability Prediction Models Descriptive Statistics Funding Source ab: In the process of author name disambiguation (AND), varying characteristics and noise of different blocks significantly impact disambiguation performance. In this paper, we propose a block‐based adaptive hyperparameter optimization method that assigns optimal hyperparameters to each block without altering the original AND model structure. Based on this, a random forest model is trained using the optimized results to fit the relationship between the block's data features and its optimal hyperparameters, thereby enabling the prediction of hyperparameters for new blocks. Empirical studies on 6 state‐of‐the‐art AND algorithms, 11 public datasets, and a manually labeled dataset of China's information and communication technology (ICT) industry patents demonstrate that the proposed method significantly outperforms the original algorithms across multiple standard performance evaluation metrics (Cluster F1/Pairwise F1, B‐Cubed F1, and K metrics). The results of the random forest regression indicate that the selected 16 features effectively predict the optimal hyperparameters. Further analysis reveals a power‐law relationship between relative block size and both relative performance and relative optimized performance across all datasets and evaluation metrics, and the relative performance improvement of the adaptive hyperparameter optimization algorithm is particularly significant for smaller blocks. These findings provide theoretical support and practical guidance for the development of AND algorithms. pubtype: Academic Journal doctype: equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|