Main Article Content
Abstract
Corporate ESG disclosure has shifted toward narrative form, yet vendor ESG ratings that most empirical work depends on show substantial disagreement across providers, and encoders dominant in financial text analysis inherit a 512-token ceiling that discards most of a typical 10-K. HET-ESGFormer combines hierarchical long-document encoding, heterogeneity-aware sparse expert routing conditioned on industry sector and firm size, and a dual-channel explanation layer subject to faithfulness verification. The framework is evaluated on 3,182 U.S.-listed firms over 2015 to 2024 using 10-K filings, sustainability reports, earnings call transcripts, and ESG-tagged news. Against nine baselines spanning econometric, dictionary-based, contextual encoder, sparse-attention, and time-series approaches, HET-ESGFormer reduces return RMSE from 0.0417 to 0.0389 and QLIKE from 0.148 to 0.128, with Diebold-Mariano tests rejecting equal predictive accuracy against every baseline after correction for multiple comparisons. The text channel's marginal contribution to downside risk prediction is materially larger than its contribution to return prediction, widening at longer forecast horizons, indicating that ESG narratives carry more information about risk than about expected return. Ablations reveal a decoupling between predictive skill and explanation faithfulness: removing the consistency regularizer barely affects accuracy but sharply degrades explanation quality. The dual-channel explanation improves ERASER comprehensiveness by 50% over SHAP applied to XGBoost, and human raters significantly prefer the framework's rationales over SHAP-based rationales on trustworthiness. Decile-sorted long-short portfolios deliver a post-cost Sharpe ratio of 1.31 against 0.74 and 0.62 for portfolios built on Refinitiv and MSCI ratings, and pass VaR backtests at 95% coverage. The results support text-based ESG modeling as an alternative to vendor-rating-based approaches, particularly where downside risk management and traceable explanation are jointly required.
Keywords
Article Details
References
- Z. Si, D. A Ali, R. B. Rosli, A. Bhaumik, and A. Ghosh, "Omni-channel retail marketing effect evaluation framework integrating big data and artificial intelligence," Edelweiss Applied Science and Technology, vol. 9, no. 3, pp. 568-583, 2025, doi: 10.55214/25768484.v9i3.5255.
- D. M. Christensen, G. Serafeim, and A. Sikochi, "Why is corporate virtue in the eye of the beholder? The case of ESG ratings," The Accounting Review, vol. 97, no. 1, pp. 147-175, 2022, doi: 10.2308/TAR-2019-0506.
- R. Gibson Brandon, P. Krueger, and P. S. Schmidt, "ESG rating disagreement and stock returns," Financial analysts journal, vol. 77, no. 4, pp. 104-127, 2021, doi: 10.1080/0015198X.2021.1963186.
- I. Beltagy, M. E. Peters, and A. Cohan, "Longformer: The long-document transformer," arXiv preprint arXiv:2004.05150, 2020.
- M. Khan, G. Serafeim, and A. Yoon, "Corporate sustainability: First evidence on materiality," The accounting review, vol. 91, no. 6, pp. 1697-1724, 2016, doi: 10.2308/accr-51383.
- J. Assael, L. Carlier, and D. Challet, "Dissecting the explanatory power of ESG features on equity returns by sector, capitalization, and year with interpretable machine learning," Journal of Risk and Financial Management, vol. 16, no. 3, p. 159, 2023, doi: 10.3390/jrfm16030159.
- F. S. Khan, S. S. Mazhar, K. Mazhar, D. A. AlSaleh, and A. Mazhar, "Model-agnostic explainable artificial intelligence methods in finance: a systematic review, recent developments, limitations, challenges and future directions," Artificial Intelligence Review, vol. 58, no. 8, p. 232, 2025, doi: 10.1007/s10462-025-11215-9.
- F. Berg, J. F. Kölbel, and R. Rigobon, "Aggregate confusion: The divergence of ESG ratings," Review of finance, vol. 26, no. 6, pp. 1315-1344, 2022, doi: 10.1093/rof/rfac033.
- Z. Si, W. Wei, W. Feng, and B. Li, "Development and Construction of Intelligent Security Monitoring System," In 2020 IEEE 3rd International Conference on Information Systems and Computer Aided Education (ICISCAE), pp. 335-338. IEEE, 2020, doi: 10.1109/ICISCAE51034.2020.9236915.
- A. H. Huang, H. Wang, and Y. Yang, "FinBERT: A large language model for extracting information from financial text," Contemporary Accounting Research, vol. 40, no. 2, pp. 806-841, 2023, doi: 10.1111/1911-3846.12832.
- S. Mehra, R. Louka, and Y. Zhang, "Esgbert: Language model to help with classification tasks related to companies environmental, social, and governance practices," arXiv preprint arXiv:2203.16788, 2022, doi: 10.5121/csit.2022.120616.
- J. A. Bingler, M. Kraus, M. Leippold, and N. Webersinke, "Cheap talk and cherry-picking: What ClimateBert has to say on corporate climate risk disclosures," Finance Research Letters, vol. 47, p. 102776, 2022, doi: 10.1016/j.frl.2022.102776.
- H.-M. Lu, Y.-T. Chien, H.-H. Yen, and Y.-H. Chen, "Utilizing Pre-Trained Language Models and Large Language Models for 10-K Items Segmentation," Journal of Information Systems, pp. 1-21, 2026, doi: 10.2308/ISYS-2025-005.
- I. Chalkidis, X. Dai, M. Fergadiotis, P. Malakasiotis, and D. Elliott, "An exploration of hierarchical attention transformers for efficient long document classification," arXiv preprint arXiv:2210.05529, 2022.
- W. Fedus, B. Zoph, and N. Shazeer, "Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity," Journal of Machine Learning Research, vol. 23, no. 120, pp. 1-39, 2022.
- R. Sassen, A.-K. Hinze, and I. Hardeck, "Impact of ESG factors on firm risk in Europe," Journal of business economics, vol. 86, no. 8, pp. 867-904, 2016, doi: 10.1007/s11573-016-0819-3.
- A. Kendall, Y. Gal, and R. Cipolla, "Multi-task learning using uncertainty to weigh losses for scene geometry and semantics," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7482-7491, doi: 10.1109/CVPR.2018.00781.
- M. Sundararajan, A. Taly, and Q. Yan, "Axiomatic attribution for deep networks," in International conference on machine learning, 2017: PMLR, pp. 3319-3328.
- S. Jain and B. C. Wallace, "Attention is not explanation," in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019, pp. 3543-3556, doi: 10.18653/v1/D19-1002.
- S. Wiegreffe and Y. Pinter, "Attention is not not explanation," in Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), 2019, pp. 11-20.
- P. Atanasova, J. G. Simonsen, C. Lioma, and I. Augenstein, "A diagnostic study of explainability techniques for text classification," in Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), 2020, pp. 3256-3274, doi: 10.18653/v1/2020.emnlp-main.263.
- J. DeYoung, S. Jain, N. F. Rajani, E. Lehman, C. Xiong, R. Socher, and B. C. Wallace, "ERASER: A benchmark to evaluate rationalized NLP models," in Proceedings of the 58th annual meeting of the association for computational linguistics, 2020, pp. 4443-4458, doi: 10.18653/v1/2020.acl-main.408.
- Z. Si, W. Wei, B. Li, and W. Feng, "Analysis of DNA Image Encryption Effect by Logistic-Sine System Combined with Fractional Chaos Stability Theory," Journal of Imaging Science & Technology, vol. 64, no. 4, pp. 1, 2020, doi: 10.2352/J.ImagingSci.Technol.2020.64.4.040413.
- A. Z. Broder, "On the resemblance and containment of documents," in Proceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No. 97TB100171), 1997: IEEE, pp. 21-29, doi: 10.1109/SEQUEN.1997.666900.
- M. Hearst, "TextTiling: segmenting text into multi-paragraph subtopic passages, in ‘Computational linguistics’, Vol. 23," ed: MIT Press, 1997.
- M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, and L. Yang, "Big bird: Transformers for longer sequences," Advances in neural information processing systems, vol. 33, pp. 17283-17297, 2020.
- D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, and Y. Wu, "Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models," in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 1280-1297, doi: 10.18653/v1/2024.acl-long.70.
- L. Duong, T. Cohn, S. Bird, and P. Cook, "Low resource dependency parsing: Cross-lingual parameter sharing in a neural network parser," in Proceedings of the 53rd annual meeting of the Association for Computational Linguistics and the 7th international joint conference on natural language processing (volume 2: short papers), 2015, pp. 845-850, doi: 10.3115/v1/P15-2139.
- R. Koenker and G. Bassett Jr, "Regression quantiles," Econometrica: journal of the Econometric Society, pp. 33-50, 1978, doi: 10.2307/1913643.
- E. F. Fama and K. R. French, "A five-factor asset pricing model," Journal of financial economics, vol. 116, no. 1, pp. 1-22, 2015, doi: 10.1016/j.jfineco.2014.10.010.
- X. Zhu, J. Li, and Y. Wang, "Are risk disclosures in financial reports informative? A text mining-based perspective," Humanities and Social Sciences Communications, vol. 11, no. 1, p. 1653, 2024, doi: 10.1057/s41599-024-04169-w.
- Z. Si, D. A Ali, R. B. Rosli, A. Bhaumik, and A. Ghosh, "Application of autonomous intelligent customer behavior prediction model based on deep learning in retail marketing strategy optimization," Edelweiss Applied Science and Technology, vol. 9, no. 3, pp. 584-598, 2025, doi: 10.55214/25768484.v9i3.5256.
- P. H. Kupiec, "Techniques for verifying the accuracy of risk measurement models," 1995, doi: 10.3905/jod.1995.407942.
- Z. Si, "Mapping the spatial-temporal evolution of imagery in Tang poetry: a computer vision and GIS-based approach," Future Digital Technologies and Artificial Intelligence, pp. 1-6, 2026, doi: 10.55670/fpll.fdtai.2.1.1.
- F. X. Diebold and R. S. Mariano, "Comparing predictive accuracy," Journal of Business & economic statistics, vol. 20, no. 1, pp. 134-144, 2002, doi: 10.1198/073500102753410444.
- A. G. Hoepner, I. Oikonomou, Z. Sautner, L. T. Starks, and X. Y. Zhou, "ESG shareholder engagement and downside risk," Review of Finance, vol. 28, no. 2, pp. 483-510, 2024, doi: 10.1093/rof/rfad034.
References
Z. Si, D. A Ali, R. B. Rosli, A. Bhaumik, and A. Ghosh, "Omni-channel retail marketing effect evaluation framework integrating big data and artificial intelligence," Edelweiss Applied Science and Technology, vol. 9, no. 3, pp. 568-583, 2025, doi: 10.55214/25768484.v9i3.5255.
D. M. Christensen, G. Serafeim, and A. Sikochi, "Why is corporate virtue in the eye of the beholder? The case of ESG ratings," The Accounting Review, vol. 97, no. 1, pp. 147-175, 2022, doi: 10.2308/TAR-2019-0506.
R. Gibson Brandon, P. Krueger, and P. S. Schmidt, "ESG rating disagreement and stock returns," Financial analysts journal, vol. 77, no. 4, pp. 104-127, 2021, doi: 10.1080/0015198X.2021.1963186.
I. Beltagy, M. E. Peters, and A. Cohan, "Longformer: The long-document transformer," arXiv preprint arXiv:2004.05150, 2020.
M. Khan, G. Serafeim, and A. Yoon, "Corporate sustainability: First evidence on materiality," The accounting review, vol. 91, no. 6, pp. 1697-1724, 2016, doi: 10.2308/accr-51383.
J. Assael, L. Carlier, and D. Challet, "Dissecting the explanatory power of ESG features on equity returns by sector, capitalization, and year with interpretable machine learning," Journal of Risk and Financial Management, vol. 16, no. 3, p. 159, 2023, doi: 10.3390/jrfm16030159.
F. S. Khan, S. S. Mazhar, K. Mazhar, D. A. AlSaleh, and A. Mazhar, "Model-agnostic explainable artificial intelligence methods in finance: a systematic review, recent developments, limitations, challenges and future directions," Artificial Intelligence Review, vol. 58, no. 8, p. 232, 2025, doi: 10.1007/s10462-025-11215-9.
F. Berg, J. F. Kölbel, and R. Rigobon, "Aggregate confusion: The divergence of ESG ratings," Review of finance, vol. 26, no. 6, pp. 1315-1344, 2022, doi: 10.1093/rof/rfac033.
Z. Si, W. Wei, W. Feng, and B. Li, "Development and Construction of Intelligent Security Monitoring System," In 2020 IEEE 3rd International Conference on Information Systems and Computer Aided Education (ICISCAE), pp. 335-338. IEEE, 2020, doi: 10.1109/ICISCAE51034.2020.9236915.
A. H. Huang, H. Wang, and Y. Yang, "FinBERT: A large language model for extracting information from financial text," Contemporary Accounting Research, vol. 40, no. 2, pp. 806-841, 2023, doi: 10.1111/1911-3846.12832.
S. Mehra, R. Louka, and Y. Zhang, "Esgbert: Language model to help with classification tasks related to companies environmental, social, and governance practices," arXiv preprint arXiv:2203.16788, 2022, doi: 10.5121/csit.2022.120616.
J. A. Bingler, M. Kraus, M. Leippold, and N. Webersinke, "Cheap talk and cherry-picking: What ClimateBert has to say on corporate climate risk disclosures," Finance Research Letters, vol. 47, p. 102776, 2022, doi: 10.1016/j.frl.2022.102776.
H.-M. Lu, Y.-T. Chien, H.-H. Yen, and Y.-H. Chen, "Utilizing Pre-Trained Language Models and Large Language Models for 10-K Items Segmentation," Journal of Information Systems, pp. 1-21, 2026, doi: 10.2308/ISYS-2025-005.
I. Chalkidis, X. Dai, M. Fergadiotis, P. Malakasiotis, and D. Elliott, "An exploration of hierarchical attention transformers for efficient long document classification," arXiv preprint arXiv:2210.05529, 2022.
W. Fedus, B. Zoph, and N. Shazeer, "Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity," Journal of Machine Learning Research, vol. 23, no. 120, pp. 1-39, 2022.
R. Sassen, A.-K. Hinze, and I. Hardeck, "Impact of ESG factors on firm risk in Europe," Journal of business economics, vol. 86, no. 8, pp. 867-904, 2016, doi: 10.1007/s11573-016-0819-3.
A. Kendall, Y. Gal, and R. Cipolla, "Multi-task learning using uncertainty to weigh losses for scene geometry and semantics," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7482-7491, doi: 10.1109/CVPR.2018.00781.
M. Sundararajan, A. Taly, and Q. Yan, "Axiomatic attribution for deep networks," in International conference on machine learning, 2017: PMLR, pp. 3319-3328.
S. Jain and B. C. Wallace, "Attention is not explanation," in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019, pp. 3543-3556, doi: 10.18653/v1/D19-1002.
S. Wiegreffe and Y. Pinter, "Attention is not not explanation," in Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), 2019, pp. 11-20.
P. Atanasova, J. G. Simonsen, C. Lioma, and I. Augenstein, "A diagnostic study of explainability techniques for text classification," in Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), 2020, pp. 3256-3274, doi: 10.18653/v1/2020.emnlp-main.263.
J. DeYoung, S. Jain, N. F. Rajani, E. Lehman, C. Xiong, R. Socher, and B. C. Wallace, "ERASER: A benchmark to evaluate rationalized NLP models," in Proceedings of the 58th annual meeting of the association for computational linguistics, 2020, pp. 4443-4458, doi: 10.18653/v1/2020.acl-main.408.
Z. Si, W. Wei, B. Li, and W. Feng, "Analysis of DNA Image Encryption Effect by Logistic-Sine System Combined with Fractional Chaos Stability Theory," Journal of Imaging Science & Technology, vol. 64, no. 4, pp. 1, 2020, doi: 10.2352/J.ImagingSci.Technol.2020.64.4.040413.
A. Z. Broder, "On the resemblance and containment of documents," in Proceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No. 97TB100171), 1997: IEEE, pp. 21-29, doi: 10.1109/SEQUEN.1997.666900.
M. Hearst, "TextTiling: segmenting text into multi-paragraph subtopic passages, in ‘Computational linguistics’, Vol. 23," ed: MIT Press, 1997.
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, and L. Yang, "Big bird: Transformers for longer sequences," Advances in neural information processing systems, vol. 33, pp. 17283-17297, 2020.
D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, and Y. Wu, "Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models," in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 1280-1297, doi: 10.18653/v1/2024.acl-long.70.
L. Duong, T. Cohn, S. Bird, and P. Cook, "Low resource dependency parsing: Cross-lingual parameter sharing in a neural network parser," in Proceedings of the 53rd annual meeting of the Association for Computational Linguistics and the 7th international joint conference on natural language processing (volume 2: short papers), 2015, pp. 845-850, doi: 10.3115/v1/P15-2139.
R. Koenker and G. Bassett Jr, "Regression quantiles," Econometrica: journal of the Econometric Society, pp. 33-50, 1978, doi: 10.2307/1913643.
E. F. Fama and K. R. French, "A five-factor asset pricing model," Journal of financial economics, vol. 116, no. 1, pp. 1-22, 2015, doi: 10.1016/j.jfineco.2014.10.010.
X. Zhu, J. Li, and Y. Wang, "Are risk disclosures in financial reports informative? A text mining-based perspective," Humanities and Social Sciences Communications, vol. 11, no. 1, p. 1653, 2024, doi: 10.1057/s41599-024-04169-w.
Z. Si, D. A Ali, R. B. Rosli, A. Bhaumik, and A. Ghosh, "Application of autonomous intelligent customer behavior prediction model based on deep learning in retail marketing strategy optimization," Edelweiss Applied Science and Technology, vol. 9, no. 3, pp. 584-598, 2025, doi: 10.55214/25768484.v9i3.5256.
P. H. Kupiec, "Techniques for verifying the accuracy of risk measurement models," 1995, doi: 10.3905/jod.1995.407942.
Z. Si, "Mapping the spatial-temporal evolution of imagery in Tang poetry: a computer vision and GIS-based approach," Future Digital Technologies and Artificial Intelligence, pp. 1-6, 2026, doi: 10.55670/fpll.fdtai.2.1.1.
F. X. Diebold and R. S. Mariano, "Comparing predictive accuracy," Journal of Business & economic statistics, vol. 20, no. 1, pp. 134-144, 2002, doi: 10.1198/073500102753410444.
A. G. Hoepner, I. Oikonomou, Z. Sautner, L. T. Starks, and X. Y. Zhou, "ESG shareholder engagement and downside risk," Review of Finance, vol. 28, no. 2, pp. 483-510, 2024, doi: 10.1093/rof/rfad034.