{"id":null,"code":"TKO_7092","name":{"valueFi":"Evaluation of Machine Learning Methods","valueEn":"Evaluation of Machine Learning Methods","valueSv":"Evaluation of Machine Learning Methods"},"credits":5.0,"minCredits":5,"maxCredits":5,"tags":[],"createdAt":1790534323011,"contentList":[{"title":{"valueFi":"Osaamistavoitteet","valueEn":"Learning outcomes","valueSv":""},"content":{"valueFi":"After learning about fundamental data analysis techniques from the prerequisite course Data analysis and knowledge discovery, this course goes deeper into techniques for building trustworthy artificial intelligence (AI) by rigorous performance estimation with modern resampling methods. The core aim is to adopt a scientific way of thinking on the evaluation design for machine learning based AI systems, namely how to answer to more specific statistical questions regarding prediction performance evaluation that go beyond the usual generalization to unseen data, such as how well the AI system works with new data given that the new data is known to differ from the already observed training data in a specific way. For example, if AI intends to carry out prediction to a geolocation known a priori to be at certain distance away from the already known measurements, one can design the resampling based performance evaluation method accordingly. In addition to the theoretical concepts of performance estimation, students learn how to implement them in practical real-world problems, including chemistry, geoinformatics and medical informatics related case studies.","valueEn":"After learning about fundamental data analysis techniques from the prerequisite course Data analysis and knowledge discovery, this course goes deeper into techniques for building trustworthy artificial intelligence (AI) by rigorous performance estimation with modern resampling methods. The core aim is to adopt a scientific way of thinking on the evaluation design for machine learning based AI systems, namely how to answer to more specific statistical questions regarding prediction performance evaluation that go beyond the usual generalization to unseen data, such as how well the AI system works with new data given that the new data is known to differ from the already observed training data in a specific way. For example, if AI intends to carry out prediction to a geolocation known a priori to be at certain distance away from the already known measurements, one can design the resampling based performance evaluation method accordingly. In addition to the theoretical concepts of performance estimation, students learn how to implement them in practical real-world problems, including chemistry, geoinformatics and medical informatics related case studies.","valueSv":"After learning about fundamental data analysis techniques from the prerequisite course Data analysis and knowledge discovery, this course goes deeper into techniques for building trustworthy artificial intelligence (AI) by rigorous performance estimation with modern resampling methods. The core aim is to adopt a scientific way of thinking on the evaluation design for machine learning based AI systems, namely how to answer to more specific statistical questions regarding prediction performance evaluation that go beyond the usual generalization to unseen data, such as how well the AI system works with new data given that the new data is known to differ from the already observed training data in a specific way. For example, if AI intends to carry out prediction to a geolocation known a priori to be at certain distance away from the already known measurements, one can design the resampling based performance evaluation method accordingly. In addition to the theoretical concepts of performance estimation, students learn how to implement them in practical real-world problems, including chemistry, geoinformatics and medical informatics related case studies."}},{"title":{"valueFi":"Sisältö","valueEn":"Content","valueSv":""},"content":{"valueFi":"Modern machine learning techniques provide invaluable tools for a wide range of application areas. However, they are pretty much useless if we can not trust that they can really carry out the tasks they are supposed to take care of or do they work at all, reflected by the old saying: \"If you cannot measure it, you cannot improve it\". This course first covers the background theory and assumptions behind standard approaches for statistical estimation of prediction performance for machine learning based artificial intelligence methods, such as independently and identically distributed (IID) sample of data, law of large numbers and its generalization for arbitrary estimators as well as cases in which the law does not hold. We then cover the basic resampling techniques, such as hold-out and cross-validation for prediction performance estimation or for model selection, and nested cross-validation for carrying out both of them simultaneously. The course also present a general framework for designing more sophisticated  performance estimands indicating how well the AI system will perform given that the training data and new data is known a priori to differ from each other in a specific way. Practical and representative examples of this types of estimands are considered from the fields of chemoinformatics, medical informatics, geoinformatics and bioinformatics. Namely, we pose questions like how well a machine learning model generalizes to new types of measurements, such as to new data from new patients in contrast to new data from the same patients already observed in the training data, and provide group or subject level hold-out or cross-validation schemes for answering such questions. With spatial cross-validation techniques, we estimate how well a geospatial model predicts to new data known a priori to be at least a certain distance away from the geographically closest measurement in the training data. Further, we introduce cross-validation techniques for answering different kinds of of specific cold start prediction settings for learning problems with pairwise data, such as prediction of drug-target or protein-protein interaction strengths, customer-product recommendations, etc. For example the following four cases, predicting the interaction strength for new drug-target pairs of which both the drug and target component are encountered in the training data as part of some other drug-target pairs, only the drug has been encountered but not the target, only the target is encountered but not the drug, or neither the drug nor the target have been encountered, are very different problems and hence require completely different prediction performance estimation schemes. All of the cases considered in the course are based on recent research results by our group or other groups with the same interests.","valueEn":"Modern machine learning techniques provide invaluable tools for a wide range of application areas. However, they are pretty much useless if we can not trust that they can really carry out the tasks they are supposed to take care of or do they work at all, reflected by the old saying: \"If you cannot measure it, you cannot improve it\". This course first covers the background theory and assumptions behind standard approaches for statistical estimation of prediction performance for machine learning based artificial intelligence methods, such as independently and identically distributed (IID) sample of data, law of large numbers and its generalization for arbitrary estimators as well as cases in which the law does not hold. We then cover the basic resampling techniques, such as hold-out and cross-validation for prediction performance estimation or for model selection, and nested cross-validation for carrying out both of them simultaneously. The course also present a general framework for designing more sophisticated  performance estimands indicating how well the AI system will perform given that the training data and new data is known a priori to differ from each other in a specific way. Practical and representative examples of this types of estimands are considered from the fields of chemoinformatics, medical informatics, geoinformatics and bioinformatics. Namely, we pose questions like how well a machine learning model generalizes to new types of measurements, such as to new data from new patients in contrast to new data from the same patients already observed in the training data, and provide group or subject level hold-out or cross-validation schemes for answering such questions. With spatial cross-validation techniques, we estimate how well a geospatial model predicts to new data known a priori to be at least a certain distance away from the geographically closest measurement in the training data. Further, we introduce cross-validation techniques for answering different kinds of of specific cold start prediction settings for learning problems with pairwise data, such as prediction of drug-target or protein-protein interaction strengths, customer-product recommendations, etc. For example the following four cases, predicting the interaction strength for new drug-target pairs of which both the drug and target component are encountered in the training data as part of some other drug-target pairs, only the drug has been encountered but not the target, only the target is encountered but not the drug, or neither the drug nor the target have been encountered, are very different problems and hence require completely different prediction performance estimation schemes. All of the cases considered in the course are based on recent research results by our group or other groups with the same interests.","valueSv":"Modern machine learning techniques provide invaluable tools for a wide range of application areas. However, they are pretty much useless if we can not trust that they can really carry out the tasks they are supposed to take care of or do they work at all, reflected by the old saying: \"If you cannot measure it, you cannot improve it\". This course first covers the background theory and assumptions behind standard approaches for statistical estimation of prediction performance for machine learning based artificial intelligence methods, such as independently and identically distributed (IID) sample of data, law of large numbers and its generalization for arbitrary estimators as well as cases in which the law does not hold. We then cover the basic resampling techniques, such as hold-out and cross-validation for prediction performance estimation or for model selection, and nested cross-validation for carrying out both of them simultaneously. The course also present a general framework for designing more sophisticated  performance estimands indicating how well the AI system will perform given that the training data and new data is known a priori to differ from each other in a specific way. Practical and representative examples of this types of estimands are considered from the fields of chemoinformatics, medical informatics, geoinformatics and bioinformatics. Namely, we pose questions like how well a machine learning model generalizes to new types of measurements, such as to new data from new patients in contrast to new data from the same patients already observed in the training data, and provide group or subject level hold-out or cross-validation schemes for answering such questions. With spatial cross-validation techniques, we estimate how well a geospatial model predicts to new data known a priori to be at least a certain distance away from the geographically closest measurement in the training data. Further, we introduce cross-validation techniques for answering different kinds of of specific cold start prediction settings for learning problems with pairwise data, such as prediction of drug-target or protein-protein interaction strengths, customer-product recommendations, etc. For example the following four cases, predicting the interaction strength for new drug-target pairs of which both the drug and target component are encountered in the training data as part of some other drug-target pairs, only the drug has been encountered but not the target, only the target is encountered but not the drug, or neither the drug nor the target have been encountered, are very different problems and hence require completely different prediction performance estimation schemes. All of the cases considered in the course are based on recent research results by our group or other groups with the same interests."}},{"title":{"valueFi":"Suoritustavat","valueEn":"Study methods","valueSv":""},"content":{"valueFi":"Lectures, Exercises, Exam","valueEn":"Lectures, Exercises, Exam","valueSv":""}},{"title":{"valueFi":"Toteutustavat","valueEn":"Course unit methods","valueSv":""},"content":{"valueFi":"","valueEn":"","valueSv":""}},{"title":{"valueFi":"Oppimateriaalit","valueEn":"Learning material","valueSv":""},"content":{"valueFi":"","valueEn":"","valueSv":""}},{"title":{"valueFi":"Lisätiedot","valueEn":"Further information","valueSv":""},"content":{"valueFi":"","valueEn":"","valueSv":""}},{"title":{"valueFi":"Kurssikirjallisuus","valueEn":"Literature","valueSv":""},"content":{"valueFi":"","valueEn":"","valueSv":""}},{"title":{"valueFi":"Esitietovaatimukset","valueEn":"Qualifications","valueSv":""},"content":{"valueFi":"Compulsory pre-requisitements: Data analysis and knowledge discovery (TKO_3103) and Statistical data analysis (TKO_7093).","valueEn":"Compulsory pre-requisitements: Data analysis and knowledge discovery (TKO_3103) and Statistical data analysis (TKO_7093).","valueSv":""}},{"title":{"valueFi":"Arviointiasteikko","valueEn":"Assessment scale","valueSv":""},"content":{"valueFi":"0-5","valueEn":"0-5","valueSv":"0-5"}},{"title":{"valueFi":"Arviointikriteerit","valueEn":"Assessment criteria","valueSv":""},"content":{"valueFi":"","valueEn":"","valueSv":""}},{"title":{"valueFi":"Arviointikriteerit 2","valueEn":"Assessment criteria 2","valueSv":""},"content":{"valueFi":"","valueEn":"","valueSv":""}},{"title":{"valueFi":"Arviointikriteerit 3","valueEn":"Assessment criteria 3","valueSv":""},"content":{"valueFi":"","valueEn":"","valueSv":""}},{"title":{"valueFi":"Arviointikriteerit 4","valueEn":"Assessment criteria 4","valueSv":""},"content":{"valueFi":"","valueEn":"","valueSv":""}},{"title":{"valueFi":"Kielet","valueEn":"Languages","valueSv":""},"content":{"valueFi":"englanti","valueEn":"English","valueSv":"engelska"}},{"title":{"valueFi":"Taso","valueEn":"Level","valueSv":""},"content":{"valueFi":"Syventävät opinnot","valueEn":"Advanced Studies","valueSv":""}},{"title":{"valueFi":"Oppiaine","valueEn":"Subject","valueSv":""},"content":{"valueFi":"Tietojenkäsittelytieteet","valueEn":"Computer Science","valueSv":"Computer Science"}},{"title":{"valueFi":"Vastuuhenkilöt","valueEn":"Person in charge","valueSv":""},"content":{"valueFi":"Tapio Pahikkala","valueEn":"Tapio Pahikkala","valueSv":"Tapio Pahikkala"}},{"title":{"valueFi":"Luokittelu","valueEn":"Classification","valueSv":""},"content":{"valueFi":"","valueEn":"","valueSv":""}},{"title":{"valueFi":"Linkit","valueEn":"Links","valueSv":""},"content":{"valueFi":"","valueEn":"","valueSv":""}}]}