Main Article Content

Abstract

Multi-table join optimization in cloud-native databases faces challenges of storage-compute separation, resource elasticity, and runtime cost variation that classic optimizers cannot handle, which drives the need for methods able to adjust their decision boundaries according to the runtime behavior of the underlying cloud infrastructure. The proposed hybrid architecture freezes the stable sub-components of an IKKBZ-GOO heuristic seed plan into a structural freeze mask that constrains the learning algorithm throughout training, rather than being discarded after a single warm-start iteration. The mask works together with a state-aware cost function whose single state-dependent parameter is updated online. Two further components complete the design: a logistic gate trained on production telemetry, and a regression containment strategy that falls back to the heuristic plan when the learned plan underperforms. Training converges in 6.4 hours on one 16-vCPU compute node without GPU acceleration. The method was evaluated on the Join Order Benchmark, TPC-H, and TPC-DS, with PostgreSQL running on an eight-node Kubernetes cluster with object-store-backed shared storage. It achieves a geometric mean speedup of 1.38-fold on JOB, 1.28-fold on TPC-DS, and 1.07-fold on TPC-H, produces a query plan in 4.6 to 6.4 ms, returns median latency to within 5% of its pre-scaling value within 22 seconds of a scaling event, and incurs a 21% overhead at the 95th percentile under node jitter. These findings demonstrate that heuristic and machine-learning components can be integrated to yield a usable optimizer whose decision boundaries adapt to the environment in which it is deployed.

Keywords

Cloud-native database Heuristic search Multi-table join optimization Query optimizer Reinforcement learning

Article Details

How to Cite
Ma, H., & Liu, S. (2026). A hybrid heuristic–reinforcement learning method for multi-table join optimization in cloud-native databases . Future Technology, 5(4), 119–132. Retrieved from https://fupubco.com/futech/article/view/1122
Bookmark and Share

References

  1. A. Verbitski, A. Gupta, D. Saha, M. Brahmadesam, K. Gupta, R. Mittal, S. Krishnamurthy, S. Maurice, T. Kharatishvili, and X. Bao, “Amazon Aurora: Design considerations for high throughput cloud-native relational databases,” in Proc. 2017 ACM Int. Conf. Management of Data (SIGMOD '17), Chicago, IL, USA, May 2017, pp. 1041–1052, doi: https://doi.org/10.1145/3035918.3056101.
  2. P. Antonopoulos, A. Budovski, C. Diaconu, A. Hernandez Saenz, J. Hu, H. Kodavalla, D. Kossmann, S. Lingam, U. F. Minhas, N. Prakash, V. Purohit, H. Qu, C. S. Ravella, K. Reisteter, S. Shrotri, D. Tang, and V. Wakade, “Socrates: The new SQL Server in the cloud,” in Proc. 2019 Int. Conf. Management of Data (SIGMOD '19), Amsterdam, The Netherlands, Jun.–Jul. 2019, pp. 1743–1756, doi: https://doi.org/10.1145/3299869.3314047.
  3. W. Cao, Y. Zhang, X. Yang, F. Li, S. Wang, Q. Hu, X. Cheng, Z. Chen, Z. Liu, J. Fang, B. Wang, Y. Wang, H. Sun, Z. Yang, Z. Cheng, S. Chen, J. Wu, W. Hu, J. Zhao, Y. Gao, S. Cai, Y. Zhang, and J. Tong, “PolarDB Serverless: A cloud native database for disaggregated data centers,” in Proc. 2021 Int. Conf. Management of Data (SIGMOD '21), Virtual Event, China, Jun. 2021, pp. 2477–2489, doi: https://doi.org/10.1145/3448016.3457560.
  4. B. Dageville, T. Cruanes, M. Zukowski, V. Antonov, A. Avanes, J. Bock, J. Claybaugh, D. Engovatov, M. Hentschel, J. Huang, A. W. Lee, A. Motivala, A. Q. Munir, S. Pelley, P. Povinec, G. Rahn, S. Triantafyllis, and P. Unterbrunner, “The Snowflake elastic data warehouse,” in Proc. 2016 Int. Conf. Management of Data (SIGMOD '16), San Francisco, CA, USA, Jun. 2016, pp. 215–226, doi: https://doi.org/10.1145/2882903.2903741.
  5. N. Armenatzoglou, S. Basu, N. Bhanoori, M. Cai, N. Chainani, K. Chinta, V. Govindaraju, T. J. Green, M. Gupta, S. Hillig, E. Hotinger, Y. Leshinksy, J. Liang, M. McCreedy, F. Nagel, I. Pandis, P. Parchas, R. Pathak, O. Polychroniou, F. Rahman, G. Saxena, G. Soundararajan, S. Subramanian, and D. Terry, “Amazon Redshift re-invented,” in Proc. 2022 Int. Conf. Management of Data (SIGMOD '22), Philadelphia, PA, USA, Jun. 2022, pp. 2205–2217, doi: https://doi.org/10.1145/3514221.3526045.
  6. R. Taft, I. Sharif, A. Matei, N. VanBenschoten, J. Lewis, T. Grieger, K. Niemi, A. Woods, A. Birzin, R. Poss, P. Bardea, A. Ranade, B. Darnell, B. Gruneir, J. Jaffray, L. Zhang, and P. Mattis, “CockroachDB: The resilient geo-distributed SQL database,” in Proc. 2020 ACM SIGMOD Int. Conf. Management of Data (SIGMOD '20), Portland, OR, USA, Jun. 2020, pp. 1493–1509, doi: https://doi.org/10.1145/3318464.3386134.
  7. Z. Si, W. Wei, B. Li, W. Feng, “Analysis of DNA Image Encryption Effect by Logistic-Sine System Combined with Fractional Chaos Stability Theory,” Journal of Imaging Science & Technology, vol. 64, no. 4, 2020, doi: https://doi.org/10.2352/J.ImagingSci.Technol.2020.64.4.040413.
  8. P. G. Selinger, M. M. Astrahan, D. D. Chamberlin, R. A. Lorie, and T. G. Price, “Access path selection in a relational database management system,” in Proc. 1979 ACM SIGMOD Int. Conf. Management of Data (SIGMOD '79), Boston, MA, USA, May 1979, pp. 23–34, doi: https://doi.org/10.1145/582095.582099.
  9. M. Steinbrunn, G. Moerkotte, and A. Kemper, “Heuristic and randomized optimization for the join ordering problem,” VLDB J., vol. 6, no. 3, pp. 191–208, Aug. 1997, doi: https://doi.org/10.1007/s007780050040.
  10. R. Marcus, P. Negi, H. Mao, C. Zhang, M. Alizadeh, T. Kraska, O. Papaemmanouil, and N. Tatbul, “Neo: A learned query optimizer,” Proc. VLDB Endow., vol. 12, no. 11, pp. 1705–1718, Jul. 2019, doi: https://doi.org/10.14778/3342263.3342644.
  11. Z. Yang, B. Chandramouli, C. Wang, J. Gehrke, Y. Li, U. F. Minhas, P.-Å. Larson, D. Kossmann, and R. Acharya, “Qd-tree: Learning data layouts for big data analytics,” in Proc. 2020 ACM SIGMOD Int. Conf. Management of Data (SIGMOD '20), Portland, OR, USA, Jun. 2020, pp. 193–208, doi: https://doi.org/10.1145/3318464.3389770.
  12. X. Yu, G. Li, C. Chai, and N. Tang, “Reinforcement learning with Tree-LSTM for join order selection,” in Proc. 2020 IEEE 36th Int. Conf. Data Engineering (ICDE), Dallas, TX, USA, Apr. 2020, pp. 1297–1308, doi: https://doi.org/10.1109/ICDE48307.2020.00116.
  13. R. Marcus, P. Negi, H. Mao, N. Tatbul, M. Alizadeh, and T. Kraska, “Bao: Making learned query optimization practical,” in Proc. 2021 Int. Conf. Management of Data (SIGMOD '21), Virtual Event, China, Jun. 2021, pp. 1275–1288, doi: https://doi.org/10.1145/3448016.3452838.
  14. Z. Si, W. Wei, W. Feng, B. Li, “Development and Construction of Intelligent Security Monitoring System,” in Proc. IEEE 3rd Int. Conf. On Information Systems and Computer Aided Education (ICISCAE), pp. 335-338, Sep. 2020, doi: https://doi.org/10.1109/ICISCAE51034.2020.9236915.
  15. S. Hasan, S. Thirumuruganathan, J. Augustine, N. Koudas, and G. Das, “Deep learning models for selectivity estimation of multi-attribute queries,” in Proc. 2020 ACM SIGMOD Int. Conf. Management of Data (SIGMOD '20), Portland, OR, USA, Jun. 2020, pp. 1035–1050, doi: https://doi.org/10.1145/3318464.3389741.
  16. B. Hilprecht, A. Schmidt, M. Kulessa, A. Molina, K. Kersting, and C. Binnig, “DeepDB: Learn from data, not from queries!,” Proc. VLDB Endow., vol. 13, no. 7, pp. 992–1005, Mar. 2020, doi: https://doi.org/10.14778/3384345.3384349.
  17. Z. Yang, W.-L. Chiang, S. Luan, G. Mittal, M. Luo, and I. Stoica, “Balsa: Learning a query optimizer without expert demonstrations,” in Proc. 2022 Int. Conf. Management of Data (SIGMOD '22), Philadelphia, PA, USA, Jun. 2022, pp. 931–944, doi: https://doi.org/10.1145/3514221.3517885.
  18. I. Trummer, J. Wang, Z. Wei, D. Maram, S. Moseley, S. Jo, J. Antonakakis, and A. Rayabhari, “SkinnerDB: Regret-bounded query evaluation via reinforcement learning,” ACM Trans. Database Syst., vol. 46, no. 3, Art. no. 9, pp. 1–45, Sep. 2021, doi: https://doi.org/10.1145/3464389.
  19. M. Ramadan, A. El-Kilany, H. M. O. Mokhtar, and I. Sobh, “RL_QOptimizer: A reinforcement learning based query optimizer,” IEEE Access, vol. 10, pp. 70502–70515, Jul. 2022, doi: https://doi.org/10.1109/ACCESS.2022.3187102.
  20. P. Negi, R. Marcus, A. Kipf, H. Mao, N. Tatbul, T. Kraska, and M. Alizadeh, “Flow-Loss: Learning cardinality estimates that matter,” Proc. VLDB Endow., vol. 14, no. 11, pp. 2019–2032, Jul. 2021, doi: https://doi.org/10.14778/3476249.3476259.
  21. V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529–533, Feb. 2015, doi: https://doi.org/10.1038/nature14236.
  22. X. Chen, H. Chen, Z. Liang, S. Liu, J. Wang, K. Zeng, H. Su, and K.-L. Tan, “LEON: A new framework for ML-aided query optimization,” Proc. VLDB Endow., vol. 16, no. 9, pp. 2261–2273, May 2023, doi: https://doi.org/10.14778/3598581.3598597.
  23. V. Leis, A. Gubichev, A. Mirchev, P. Boncz, A. Kemper, and T. Neumann, “How good are query optimizers, really?,” Proc. VLDB Endow., vol. 9, no. 3, pp. 204–215, Nov. 2015, doi: https://doi.org/10.14778/2850583.2850594.
  24. Z. Yang, A. Kamsetty, S. Luan, E. Liang, Y. Duan, X. Chen, and I. Stoica, “NeuroCard: One cardinality estimator for all tables,” Proc. VLDB Endow., vol. 14, no. 1, pp. 61–73, Sep. 2020, doi: https://doi.org/10.14778/3421424.3421432.
  25. R. Zhu, W. Chen, B. Ding, X. Chen, A. Pfadler, Z. Wu, and J. Zhou, “Lero: A learning-to-rank query optimizer,” Proc. VLDB Endow., vol. 16, no. 6, pp. 1466–1479, Feb. 2023, doi: https://doi.org/10.14778/3583140.3583160.
  26. X. Yu, C. Chai, G. Li, and J. Liu, “Cost-based or learning-based? A hybrid query optimizer for query plan selection,” Proc. VLDB Endow., vol. 15, no. 13, pp. 3924–3936, Sep. 2022, doi: https://doi.org/10.14778/3565838.3565846.
  27. J. Chen, G. Ye, Y. Zhao, S. Liu, L. Deng, X. Chen, R. Zhou, and K. Zheng, “Efficient join order selection learning with graph-based representation,” in Proc. 28th ACM SIGKDD Conf. Knowledge Discovery and Data Mining (KDD '22), Washington, DC, USA, Aug. 2022, pp. 97–107, doi: https://doi.org/10.1145/3534678.3539303.
  28. L. Weng, R. Zhu, D. Wu, B. Ding, B. Zheng, and J. Zhou, “Eraser: Eliminating performance regression on learned query optimizer,” Proc. VLDB Endow., vol. 17, no. 5, pp. 926–938, Jan. 2024, doi: https://doi.org/10.14778/3641204.3641205.
  29. Y. Han, Z. Wu, P. Wu, R. Zhu, J. Yang, L. W. Tan, K. Zeng, G. Cong, Y. Qin, A. Pfadler, Z. Qian, J. Zhou, J. Li, and B. Cui, “Cardinality estimation in DBMS: A comprehensive benchmark evaluation,” Proc. VLDB Endow., vol. 15, no. 4, pp. 752–765, Dec. 2021, doi: https://doi.org/10.14778/3503585.3503586.
  30. P. Wu and Z. G. Ives, “Modeling shifting workloads for learned database systems,” Proc. ACM Manag. Data, vol. 2, no. 1, Art. no. 38, pp. 1–27, Mar. 2024, doi: https://doi.org/10.1145/3639293.