Please use this identifier to cite or link to this item: https://hdl.handle.net/11147/3447
Full metadata record
DC FieldValueLanguage
dc.contributor.advisorPüskülcü, Halisen
dc.contributor.authorŞimşek, Kadir-
dc.date.accessioned2014-07-22T13:51:33Z-
dc.date.available2014-07-22T13:51:33Z-
dc.date.issued2004en
dc.identifier.urihttp://hdl.handle.net/11147/3447-
dc.descriptionThesis (Master)--Izmir Institute of Technology, Computer Engineering, Izmir, 2004en
dc.descriptionIncludes bibliographical references (leaves: 61-63)en
dc.descriptionText in English; Abstract: Turkish and Englishen
dc.descriptionix, 70 leavesen
dc.description.abstractIn this study of topic .Categorization of Web Sites in Turkey with SVM. after a brief introduction to what the World Wide Web is and a more detailed description of text categorization and web site categorization concepts, categorization of web sites including all prerequisites for classification task takes part. As an information resource the web has an undeniable importance in human life. However the huge structure of the web and its uncontrolled growth led to new information retrieval research areas to be risen in last years. Web mining, the general name of these studies, investigates activities and structures on the web to automatically discover and gather meaningful information from the web documents. It consists of three subfields: .Web Structure Mining., .Web Content Mining. and .Web Usage Mining.. In this project, web content mining concept was applied on the web sites in Turkey during the categorization process. Support Vector Machine, a supervised learning method based on statistics and principle of structural risk minimization is used as the machine learning technique for web site categorization. This thesis is intended to draw a conclusion about web site distributions with respect to thematic categorization based on text. The popular web directory Yahoo.s 12 top level categories were used in this project. Beside of the main purpose, we gathered several statistical descriptive informations about web sites and contents used in html pages. Metatag usage percentages, html design structures and plug-in usage are some of these information. The processes taken through solution, start with employing a web downloader which downloads web page contents and other information such as frame content from each web site. Next, manipulating, parsing and simplifying the downloaded documents takes place. At this point, preperations for categorization task are completed. Then, by applying Support Vector Machine (SVM) package SVMLight developed by Thorsten Joachims, web sites are classified under given categories. The classification results obtained in the last section show that there are some over-lapping categories exist and accuracy and precision values are between 60-80. In addition to categorization results, we saw that almost 17 of web sites utilize html frames and 9367 web sites include metakeywords.en
dc.language.isoenen_US
dc.publisherIzmir Institute of Technologyen
dc.rightsinfo:eu-repo/semantics/openAccessen_US
dc.subject.lccQA76.9.D343 .S58 2004en
dc.subject.lcshData miningen
dc.subject.lcshWeb sites--Turkeyen
dc.subject.lcshWeb sites--Classificationen
dc.titleCategorization of web sites in Turkey with SVMen_US
dc.typeMaster Thesisen_US
dc.institutionauthorŞimşek, Kadir-
dc.departmentThesis (Master)--İzmir Institute of Technology, Computer Engineeringen_US
dc.relation.publicationcategoryTezen_US
item.fulltextWith Fulltext-
item.grantfulltextopen-
item.languageiso639-1en-
item.openairecristypehttp://purl.org/coar/resource_type/c_18cf-
item.cerifentitytypePublications-
item.openairetypeMaster Thesis-
Appears in Collections:Master Degree / Yüksek Lisans Tezleri
Files in This Item:
File Description SizeFormat 
T000450.pdfMasterThesis960.12 kBAdobe PDFThumbnail
View/Open
Show simple item record



CORE Recommender

Page view(s)

168
checked on Nov 18, 2024

Download(s)

118
checked on Nov 18, 2024

Google ScholarTM

Check





Items in GCRIS Repository are protected by copyright, with all rights reserved, unless otherwise indicated.