Automated Feature Engineering Techniques for Tabular Data

  • Authors

    • Prof. Yuki Tanaka Professor of Intelligence Information Science, Data Science, Kyoto University, Japan. Author
    • Kenji Sato Engineering Director, Sony Corporation, Japan. Author

    DOI:

    https://doi.org/10.67228/30713315/IJAIDT-2023PI7W59

    Published 05-05-2023

  • Automated Feature Engineering, Tabular Data, Machine Learning, AutoML, Feature Selection, Data Preprocessing

    Issue

    Section

    Articles

    How to Cite

    [1]
    Y. Tanaka and K. Sato, “Automated Feature Engineering Techniques for Tabular Data”, IJAIDT, vol. 6, no. 1, pp. 01–15, May 2023, doi: 10.67228/30713315/IJAIDT-2023PI7W59.
  • Abstract

    AFE is an automated feature engineering system that has become a key enabler to scalable and high-performing machine learning systems with scalable systems that run on tabular data. The conventional feature engineering makes excessive use of domain, trial and error and iterative optimization, which are computed consuming time, error prone and hard to repeat. As data-driven applications in finance, healthcare, manufacturing, and e-commerce have been exponentially increasing, there is an increasing need in automated, systematic, and reliable methods of features construction. The overall objective of automated Feature Engineering methods is to generate, transform, select, and optimize features adventurously, from raw tabular data, with minimal human intervention, and at a higher predictive efficiency. The paper contains a complete detailed analysis of automated feature engineering approaches to tabular data with references to their theoretical principles, algorithmic approaches, and real-world examples. The paper expounds on rule construction based feature construction, statistical construction, deep learning based representation, evolutionary learning, and reinforcing learning methods, and end to end AutoML. An intricate literature review shows major achievements, comparisons, and unresolved issues. The suggested methodology defines the three branches of feature generation, selection and evaluation as a single automated pipeline via mathematical representation and algorithmic processes. The effectiveness of automated feature engineering is proved by experimental results that show that the method can enhance the accuracy and robustness of models as well as improve generalization with respect to the multiple benchmark datasets. Lastly, issues of limitations, interpretability, computational trade-offs, and research directions are discussed in the paper. The given publication meets the IEEE publication standards and offers well-organized, high-quality information to a researcher or an organization practitioner dealing with tabular data analytics.

  • References

    [1] Guyon, I., & Elisseeff, A. (2003). An introduction to variable and feature selection. Journal of Machine Learning Research, 3, 1157–1182.

    [2] Liu, H., & Motoda, H. (1998). Feature selection for knowledge discovery and data mining. Springer.

    [3] Kuhn, M., & Johnson, K. (2013). Applied predictive modeling. Springer.

    [4] Kotsiantis, S., Kanellopoulos, D., & Pintelas, P. (2006). Data preprocessing for supervised learning. International Journal of Computer Science, 1(2), 111–117.

    [5] Jensen, D., & Neville, J. (2002). Linkage and autocorrelation cause feature selection bias in relational learning. ICML.

    [6] Deepayan, C., & Getoor, L. (2006). Link-based classification. SIGKDD Explorations, 8(2), 14–23.

    [7] Kanter, J. M., & Veeramachaneni, K. (2015). Deep feature synthesis: Towards automating data science endeavors. IEEE DSAA.

    [8] Markovitch, S., & Rosenstein, D. (2002). Feature generation using general constructor functions. Machine Learning, 49(1), 59–98.

    [9] Koza, J. R. (1992). Genetic programming: On the programming of computers by means of natural selection. MIT Press.

    [10] Neshatian, K., Zhang, M., & Andreae, P. (2012). A survey on evolutionary computation approaches to feature selection and construction. IEEE Transactions on Evolutionary Computation, 16(6), 817–840.

    [11] Bengio, Y., Courville, A., & Vincent, P. (2013). Representation learning: A review and new perspectives. IEEE TPAMI, 35(8), 1798–1828.

    [12] Hinton, G. E., & Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks. Science, 313(5786), 504–507.

    [13] Guo, C., Berkhahn, F. (2016). Entity embeddings of categorical variables. arXiv preprint arXiv:1604.06737.

    [14] He, X., Zhao, K., & Chu, X. (2021). AutoML: A survey of the state-of-the-art. Knowledge-Based Systems, 212, 106622.

    [15] Feurer, M., et al. (2015). Efficient and robust automated machine learning. NeurIPS, 2962–2970.

  • Downloads