Some datasets were taken from the UCI repository (Merz, C. J., and Murphy, P. M. (1996). UCI Repository of machine learning databases, Irvine, CA: University of California, Department of Information and Computer Science) found at: http://www.ics.uci.edu/~mlearn/MLRepository.html .
Several papers describing the technology used in MineSet are available at: http://www.sgi.com/software/mineset/mineset_data.html .
An excellent, non-technical introduction to data mining techniques is:
Michael Berry and Gordon Linoff. Data Mining Techniques. New York: John Wiley & Sons, 1997. ISBN 0-471-17980-9. See http://www.sgi.com/software/mineset/mineset_data.html http://www.data-miners.com/ .
A comparative study of data mining tools, including MineSet, was done by the Two Crows Corporation. It contains a good introduction to data mining.
Two Crows Corporation. Data Mining: Products, Applications & Technologies. Ordering information is available at http://www.twocrows.com .
A paper describing MLC++, the underlying analytical engine used in MineSet, is described in:
Kohavi, R., Sommerfield, D., Dougherty, J., Data Mining using MLC++, a Machine Learning Library in C++. International Journal of Artificial Intelligence Tools, Vol. 6, No. 4, 1997, p. 537-566. See http://robotics.stanford.edu/users/ronnyk/ .
A general and easy-to-read introduction to machine learning is:
Weiss, S. M., and C. A. Kulikowski. Computer Systems that Learn. San Mateo, CA: Morgan Kaufmann Publishers, Inc., 1991.
A general comparison of algorithms and descriptions is provided in:
Taylor, C., D. Michie, and D. Spiegalhalter. Machine Learning, Neural and Statistical Classification. Paramount Publishing International, 1994.
An easy-to-read introduction to decision tree induction is:
Quinlan, J. R. C4.5: Programs for Machine Learning. Los Altos, CA: Morgan Kaufmann Publishers, Inc., 1993.
An excellent book on decision trees from a statistical perspective is:
Breiman, L., J. H. Friedman, R. A. Olshen, and C.J. Stone. Classification and Regression Trees. Wadsworth International Group, 1984.
A good edited volume of machine learning techniques is:
Dietterich, T. G. and J. W. Shavlik (Eds). Readings in Machine Learning. Morgan Kaufmann Publishers, Inc., 1990.
A summary of accuracy estimation techniques is given in:
Kohavi, R. A study of cross-validation and bootstrap for accuracy estimation and model selection. In Proceedings of the 14th International Joint Conference on Artificial Intelligence, edited by C. S. Mellish. Morgan Kaufmann Publishers, Inc., 1995. Available at http://robotics.stanford.edu/users/ronnyk/ .
The following papers describes the decision table metaphor:
Kohavi, R., The Power of Decision Tables. In The European Conference on Machine Learning, 1995.
Becker, B. Visualizing Decision Tables. In IEEEs Proceedings of Information Visualization, 1998.
A good reference to a paper explaining that no classifier can be “best” is:
Schaffer, C. A conservation law for generalization performance. In Machine Learning: Proceedings of the Eleventh International Conference, 259-265. Morgan Kaufmann Publishers, Inc., 1994. Available at
http://wwwcs.hunter.cuny.edu/faculty/schaffer/papers/list.html
.
MineSet uses an advanced version of the Option Trees described in:
Ron Kohavi and Clayton Kunz. Option Decision Trees with Majority Votes. Machine Learning: Proceedings of the Fourteenth International Conference”, Morgan Kaufmann Publishers, Inc., 1997. (See http://robotics.stanford.edu/users/ronnyk ). The option trees used in MineSet average the predictions and do not simply vote them as described in this paper. Option Trees were first introduced by Wray Buntine in his thesis A Theory of Learning Classification Rules, 1992, School of Computing Science, University of Technology, Sydney.
The following paper describes the wrapper method used to select the features for the Evidence Classifier:
Kohavi, R., Sommerfield, D. (1995). Feature Subset Selection Using the Wrapper Model: Overfitting and Dynamic Search Space Topology. The First International Conference on Knowledge Discovery and Data Mining, pp. 192-197. Available at:
http://robotics.stanford.edu/users/ronnyk/
.
An excellent introduction to the Evidence Classifier (Naive-Bayes) is:
Kononenko, I. (1993). Inductive and Bayesian Learning in Medical Diagnosis. Applied Artificial Intelligence, pp. 7:317-337.
The following paper describes conditions under which the Evidence Inducer is optimal:
Domingos, P., Pazzani, M., Beyond Independence: Conditions for the Optimality of the Simple Bayes Classifier. Machine Learning, Volume 29, No. 2/3, Nov/Dec 1997, pp. 103-130.
The following paper describes the use of the wrapper method in the Evidence Inducer:
Kohavi, R., John, G., Wrappers for Feature Subset Selection. In Artificial Intelligence Journal, special issue on relevance, Vol. 97, Nos 1-2, pp. 273-324.
The following paper describes the Laplace correction option:
Cestnik, B. (1990). Estimating Probabilities: A Crucial Task in Machine Learning. Proceedings of the Ninth European Conference on Artificial Intelligence, pp. 147-149.
The following paper describes the automatic Laplace correction used in MineSet:
Kohavi R., Becker B., and Sommerfield D., Improving Simple Bayes, European Conference on Machine Learning, 1997 (poster). Available at: http://robotics.stanford.edu/users/ronnyk/ .
The following paper describes the Evidence Classifier (Naive-Bayes):
Langley, P., Iba, W., Thompson, K. (1992). An Analysis of Bayesian Classifiers. Proceedings of the Tenth National Conference on Artificial Intelligence, pp. 223-228. Available at: http://www.isle.org/~langley/pubs.html .
Simple Bayesian Classifier. To appear in Lecture Notes in Computer Science: Issues in the Integration of Data Mining and Data Visualization, Springer Verlag, 1998.
The following books describe the Evidence Classifier:
Good, I. J. The Estimation of Probabilities: An Essay on Modern Bayesian Methods. MIT Press, 1965.
Duda, R., Hart, P. Pattern Classification and Scene Analysis, Wiley, 1973.
The following paper shows that while the conditional independence assumption can be violated, the classification accuracy of the evidence classifier (called Simple Bayes in this paper) can be good:
Domingos P., Pazzani M (1996). Beyond Independence: Conditions for the Optimality of the Simple Bayesian Classifier. Machine learning, Proceedings of the 13th International Conference (ICML '96), pp. 105-112. Available at http://www.ics.uci.edu/~pedrod/ .
The following paper describes and provides further references for the technical details of the Splat Visualizer.
Becker, Barry G, Volume Rendering for Relational Data, to appear in Proceedings of Information Visualization '97, IEEE Computer Society Press, Los Alamitos CA, October 19-24, 1997.
The following paper explains how to use Gaussian splats for volume rendering.
Westover, Lee, Footprint Evaluation for Volume Rendering in Proceedings of SIGGRAPH `90, Vol. 24, No. 4, pages 367-376).
The iris database was originally used in Fisher, R. A. 1936. The use of multiple measurements in taxonomic problems. Annals of Eugenics 7(1):179-188. It is a classical problem in many statistical texts.
The breast cancer database was obtained from Dr. William H. Wolberg, L. Mangasarian, and W. H. Wolberg. Cancer diagnosis via linear programming. SIAM News 23(5):1 & 18. University of Wisconsin Hospitals, Madison, September 1990.
The data for the mushroom sample file comes from: Audubon Society Field Guide to North American Mushrooms. New York: Alfred A. Knopf, 1981.
The data on congressional voting was taken from the Congressional Quarterly Almanac, 98th Congress, 2nd session 1984, Volume XL, Congressional Quarterly Inc.: Washington, D.C., 1985.
The adult dataset was derived from the US Census Bureau survey in 1994 (http://www.census.gov/ftp/pub/DES/www/welcome.html ).