[go: up one dir, main page]
More Web Proxy on the site http://driver.im/
Skip to main content

Constraint Based Induction of Multi-objective Regression Trees

  • Conference paper
Knowledge Discovery in Inductive Databases (KDID 2005)

Part of the book series: Lecture Notes in Computer Science ((LNISA,volume 3933))

Included in the following conference series:

  • 530 Accesses

Abstract

Constrained based inductive systems are a key component of inductive databases and responsible for building the models that satisfy the constraints in the inductive queries. In this paper, we propose a constraint based system for building multi-objective regression trees. A multi-objective regression tree is a decision tree capable of predicting several numeric variables at once. We focus on size and accuracy constraints. By either specifying maximum size or minimum accuracy, the user can trade-off size (and thus interpretability) for accuracy. Our approach is to first build a large tree based on the training data and to prune it in a second step to satisfy the user constraints. This has the advantage that the tree can be stored in the inductive database and used for answering inductive queries with different constraints. Besides size and accuracy constraints, we also briefly discuss syntactic constraints. We evaluate our system on a number of real world data sets and measure the size versus accuracy trade-off.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

Unable to display preview. Download preview PDF.

Similar content being viewed by others

References

  1. Almuallim, H.: An efficient algorithm for optimal pruning of decision trees. Artificial Intelligence 83(2), 347–362 (1996)

    Article  Google Scholar 

  2. Blockeel, H., De Raedt, L., Ramon, J.: Top-down induction of clustering trees. In: Proceedings of the 15th International Conference on Machine Learning, pp. 55–63 (1998)

    Google Scholar 

  3. Bohanec, M., Bratko, I.: Trading accuracy for simplicity in decision trees. Machine Learning 15(3), 223–250 (1994)

    MATH  Google Scholar 

  4. Breiman, L.: Bagging predictors. Machine Learning 24(2), 123–140 (1996)

    MATH  Google Scholar 

  5. Breiman, L.: Random forests. Machine Learning 45(1), 5–32 (2001)

    Article  MATH  Google Scholar 

  6. Breiman, L., Friedman, J.H., Olshen, R.A., Stone, C.J.: Classification and Regression Trees. Wadsworth, Belmont (1984)

    MATH  Google Scholar 

  7. De Raedt, L.: A perspective on inductive databases. SIGKDD Explorations 4(2), 69–77 (2002)

    Article  MathSciNet  Google Scholar 

  8. Demšar, D., Debeljak, M., Lavigne, C., Džeroski, S.: Modelling pollen dispersal of genetically modified oilseed rape within the field. Abstract presented at The Annual Meeting of the Ecological Society of America, Montreal, Canada, August 7-12 (2005)

    Google Scholar 

  9. Demšar, D., Džeroski, S., Henning Krogh, P., Larsen, T., Struyf, J.: Using multiobjective classification to model communities of soil microarthropods. Ecological Modelling (2005) (to appear)

    Google Scholar 

  10. Džeroski, S., Demšar, D., Grbović, J.: Predicting chemical parameters of river water quality from bioindicator data. Applied Intelligence 13(1), 7–17 (2000)

    Article  Google Scholar 

  11. Džeroski, S., Colbach, N., Messean, A.: Analysing the effect of field characteristics on gene flow between oilseed rape varieties and volunteers with regression trees. Submitted to the The Second International Conference on Co-existence between GM and non-GM based agricultural supply chains (GMCC 2005), Montpellier, France, November 14-15 (2005)

    Google Scholar 

  12. Garofalakis, M., Hyun, D., Rastogi, R., Shim, K.: Building decision trees with constraints. Data Mining and Knowledge Discovery 7(2), 187–214 (2003)

    Article  MathSciNet  Google Scholar 

  13. Imielinski, T., Mannila, H.: A database perspective on knowledge discovery. Communications of the ACM 39(11), 58–64 (1996)

    Article  Google Scholar 

  14. Kampichler, C., Džeroski, S., Wieland, R.: The application of machine learning techniques to the analysis of soil ecological data bases: Relationships between habitat features and collembola community characteristics. Soil Biology and Biochemistry 32, 197–209 (2000)

    Article  Google Scholar 

  15. Quinlan, J.R.: C4.5: Programs for Machine Learning. Morgan Kaufmann series in Machine Learning. Morgan Kaufmann, San Francisco (1993)

    Google Scholar 

  16. Sain, R.S., Carmack, P.S.: Boosting multi-objective regression trees. Computing Science and Statistics 34, 232–241 (2002)

    Google Scholar 

  17. Witten, I., Frank, E.: Data Mining: Practical machine learning tools and techniques, 2nd edn. Morgan Kaufmann, San Francisco (2005)

    MATH  Google Scholar 

Download references

Author information

Authors and Affiliations

Authors

Editor information

Editors and Affiliations

Rights and permissions

Reprints and permissions

Copyright information

© 2006 Springer-Verlag Berlin Heidelberg

About this paper

Cite this paper

Struyf, J., Džeroski, S. (2006). Constraint Based Induction of Multi-objective Regression Trees. In: Bonchi, F., Boulicaut, JF. (eds) Knowledge Discovery in Inductive Databases. KDID 2005. Lecture Notes in Computer Science, vol 3933. Springer, Berlin, Heidelberg. https://doi.org/10.1007/11733492_13

Download citation

  • DOI: https://doi.org/10.1007/11733492_13

  • Publisher Name: Springer, Berlin, Heidelberg

  • Print ISBN: 978-3-540-33292-3

  • Online ISBN: 978-3-540-33293-0

  • eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics