<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns="http://www.w3.org/2005/Atom">
<title>Department of  Statistics</title>
<link href="https://sure.su.ac.th/xmlui/handle/123456789/945" rel="alternate"/>
<subtitle>ภาควิชาสถิติ / สาขาวิชาสถิติประยุกต์</subtitle>
<id>https://sure.su.ac.th/xmlui/handle/123456789/945</id>
<updated>2026-09-13T14:11:09Z</updated>
<dc:date>2026-09-13T14:11:09Z</dc:date>
<entry>
<title>DIABETES CLASSIFICATION USING  MACHINE LEARNING TECHNIQUES</title>
<link href="https://sure.su.ac.th/xmlui/handle/123456789/28440" rel="alternate"/>
<author>
<name>เมธาพร ผ่องยิ่ง</name>
</author>
<id>https://sure.su.ac.th/xmlui/handle/123456789/28440</id>
<updated>2024-05-15T20:10:17Z</updated>
<published>0004-01-01T00:00:00Z</published>
<summary type="text">DIABETES CLASSIFICATION USING  MACHINE LEARNING TECHNIQUES; การจำแนกการเป็นโรคเบาหวานโดยใช้เทคนิค Machine learning
เมธาพร ผ่องยิ่ง
Nowadays, Machine learning techniques play an increasingly prominent role in medical diagnosis because using these techniques can be analyzed to find patterns or facts that are difficult to explain, which contributes to making the diagnosis more accurate. The purpose of this research is to compare the efficiency of diabetic classification models with and without interaction using four machine learning techniques including Decision tree, Random forest, Support Vector Machine and K-Nearest neighbor. These models are compared base on accuracy, precision, recall, and F1-score. The results of this research showed that the models with interaction have better classification performance than those without interaction for all 4 machine learning techniques. Among models with interaction, Random forest classifiers had the best performance with 97.5% accuracy, 97.4% precision, 96.6% recall, and 97% F1-score. In the same way, Random forest also had the best classification performance among models without interaction with  88.2% accuracy, 92.2% precision, 89.3% recall, and 90.7% F1-score. The findings from this research can be further developed into a program to effectively screen diabetes patients.; ปัจจุบันเทคนิค Machine learning ได้เข้ามามีบทบาททางการแพทย์ในการวินิจฉัยโรคมากขึ้น เนื่องจากเราสามารถใช้เทคนิค Machine learning ในการวิเคราะห์ข้อมูลขนาดใหญ่ทางการแพทย์ เพื่อค้นหารูปแบบหรือข้อเท็จจริงบางอย่างที่ยากต่อการอธิบาย ซึ่งมีส่วนช่วยให้การวินิจฉัยโรคทำได้อย่างแม่นยำมากยิ่งขึ้น โดยในงานวิจัยนี้มีวัตถุประสงค์เพื่อเปรียบเทียบประสิทธิภาพของเทคนิคที่ใช้ในการสร้างแบบจําลอง Machine learning สำหรับการจำแนกการเป็นโรคเบาหวานกรณีที่พิจารณาและไม่พิจารณาอิทธิพลร่วม 4 เทคนิค ได้แก่ เทคนิคต้นไม้ตัดสินใจ (Decision tree) เทคนิคต้นไม้ป่าสุ่ม (Random Forest) เทคนิคซัพพอร์ตเวกเตอร์แมชชีน (Support Vector Machine) และเทคนิคเพื่อนบ้านใกล้ที่สุด (K-Nearest Neighbor) โดยมีเกณฑ์ที่ใช้ในการทดสอบประสิทธิภาพของการจำแนก คือ ค่าความถูกต้อง (accuracy) ค่าความเที่ยง (precision) ค่าความครบถ้วน (recall) และค่าคะแนน F1 (F1-score) ที่ให้ค่ามากที่สุด ซึ่งผลของการวิจัยพบว่าแบบจำลองกรณีที่พิจารณาอิทธิพลร่วมมีประสิทธิภาพการจำแนกดีกว่าแบบจำลองกรณีที่ไม่พิจารณาอิทธิพลร่วมทั้ง 4 เทคนิค โดยที่แบบจำลองกรณีที่พิจารณาอิทธิพลร่วม เทคนิค Random forest มีประสิทธิภาพการจำแนกดีที่สุด ซึ่งให้ค่าความถูกต้องในการจำแนก 97.5% มีค่าความแม่นยำที่ 97.4% มีค่าความครบถ้วนที่ 96.6% และค่าคะแนน F1 ที่ 97%  ในทางเดียวกัน แบบจำลองกรณีที่ไม่พิจารณาอิทธิพลร่วม เทคนิค Random forest มีประสิทธิภาพการจำแนกดีที่สุด ซึ่งให้ค่าความถูกต้องในการจำแนก 88.2% มีค่าความแม่นยำที่ 92.2% มีค่าความครบถ้วนที่ 89.3% และค่าคะแนน F1 ที่ 90.7% โดยผลการวิจัยที่ได้นี้สามารถนำไปใช้เป็นแนวทางในการพัฒนาโปรแกรมสำหรับการคัดกรองผู้ป่วยโรคเบาหวานได้อย่างมีประสิทธิภาพต่อไปในอนาคต
</summary>
<dc:date>0004-01-01T00:00:00Z</dc:date>
</entry>
<entry>
<title>A comparison of the efficiency of the HEWMA EEWMA TEWMA and DMEWMA control charts for the right skew distributions</title>
<link href="https://sure.su.ac.th/xmlui/handle/123456789/28438" rel="alternate"/>
<author>
<name>อรณิช ชุมภูวร</name>
</author>
<id>https://sure.su.ac.th/xmlui/handle/123456789/28438</id>
<updated>2024-05-15T20:10:12Z</updated>
<published>0004-01-01T00:00:00Z</published>
<summary type="text">A comparison of the efficiency of the HEWMA EEWMA TEWMA and DMEWMA control charts for the right skew distributions; การเปรียบเทียบประสิทธิภาพของแผนภูมิควบคุม HEWMA EEWMA TEWMA และ DMEWMA สำหรับข้อมูลที่มีการแจกแจงแบบเบ้ขวา
อรณิช ชุมภูวร
The purpose of this research is to compare the efficiency detection of process parameter shift for gamma distribution and log-normal distribution Control Chart, HEWMA EEWMA TEWMA and DMEWMA. The process of this research is imitated by using Monte Carlo Simulation Technique for 10,000 iterations. The data is defined by gamma distribution with scale and location parameter those are G(1,1), G(4,1), G(2,2), G(1,2)   and defined by log-normal distribution with scale and location parameter those are ln(1,1), ln(4,1), ln(2,2), ln(1,2) . The first smoothing parameter are 0.25,0.5 and second are 0.2,0.07 for EEWMA control chart ,The first smoothing parameter are 0.1,0.3 and second 0.11,0.31 for HEWMA control chart, Smoothing parameter are 0.05,0.1 for TEWMA control chart, Smoothing parameter are 0.05,0.75 for DMEWMA control chart, and process shift sizes are 0.01 - 2 respectively. The criterion is considered by out-of-control process (Average Run Length: ARL) which is the most efficiency control chart will show the least average run length for out-of control process. The situations given above in all situations. It was found that TEWMA control charts and DMEWMA control charts tended to be more effective in detecting small changes parameters shift in process than HEWMA and EEWMA control charts. The larger the parameters shift in the range 1 to 2, the HEWMA control chart is more effective at detecting changes.; งานวิจัยนี้มีวัตถุประสงค์เพื่อเปรียบเทียบประสิทธิภาพของแผนภูมิควบคุมในการตรวจจับการเปลี่ยนแปลงขนาดเล็กของพารามิเตอร์ เมื่อกระบวนการผลิตมีการแจกแจงแกมมา และการแจกแจงล็อกนอร์มัล ได้แก่แผนภูมิควบคุม HEWMA EEWMA TEWMA และ DMEWMA  ดำเนินการโดยการจำลองข้อมูลด้วยวิธีมอลติคาร์ที่ทำซ้ำ 10,000 รอบ กำหนดให้ข้อมูลมีการแจกแจงแกมมา และการแจกแจงล็อกนอร์มัล เมื่อกระบวนการอยู่ในการควบคุมโดยกำหนดค่าพารามิเตอร์ตามสถานการณ์ที่กำหนด G(1,1), G(4,1), G(2,2), G(1,2) และ ln(1,1), ln(4,1), ln(2,2), ln(1,2)   และกำหนดค่าพารามิเตอร์ปรับให้เรียบของแผนภูมิควบคุมทั้ง 4 ชนิด ได้แก่ ค่าพารามิเตอร์ปรับเรียบตัวแรกเท่ากับ 0.25, 0.5 ตัวที่สองเท่ากับ 0.2, 0.07  สำหรับแผนภูมิควบคุม EEWMA ค่าพารามิเตอร์ปรับเรียบตัวแรกเท่ากับ 0.1, 0.3 ตัวที่สองเท่ากับ0.11, 0.31 สำหรับแผนภูมิควบคุม HEWMA ค่าพารามิเตอร์ปรับเรียบเท่ากับ 0.05, 0.1 สำหรับแผนภูมิควบคุม TEWMA และค่าพารามิเตอร์ปรับเรียบเท่ากับ 0.05, 0.75 สำหรับแผนภูมิควบคุม DMEWMA  และขนาดการเปลี่ยนแปลงกระบวนการเท่ากับ 0.01 ถึง 2โดยเกณฑ์ที่ใช้วัดประสิทธิภาพของแผนภูมิควบคุม จะพิจารณาจากค่าความยาวรันเฉลี่ยซึ่งแผนภูมิที่มีประสิทธิภาพมากที่สุดจะให้ค่าความยาวรันเฉลี่ยเมื่อกระบวนการออกนอกการควบคุมมีค่าน้อยที่สุด  ภายใต้สถานการณ์ที่กำหนดข้างต้นในทุกสถานการณ์ พบว่าแผนภูมิควบคุม TEWMA และแผนภูมิควบคุม DMEWMA มีประแนวโน้มมีประสิทธิภาพในการตรวจจับการเปลี่ยนแปลงขนาดเล็กของพารามิเตอร์ในกระบวนการผลิตได้ดีกว่าแผนภูมิควบคุม HEWMA และ EEWMA แต่บางสถานการณ์เมื่อขนาดการเปลี่ยนแปลงของกระบวนการมีค่ามากขึ้นอยู่ในช่วง 1 ถึง 2 แผนภูมิควบคุม HEWMA มีประสิทธิภาพในการตรวจจับการเปลี่ยนแปลงได้ดีกว่า
</summary>
<dc:date>0004-01-01T00:00:00Z</dc:date>
</entry>
<entry>
<title>Regularization in High Dimensional Logistic Regression Model by using Adaptive LASSO Method</title>
<link href="https://sure.su.ac.th/xmlui/handle/123456789/26952" rel="alternate"/>
<author>
<name/>
</author>
<id>https://sure.su.ac.th/xmlui/handle/123456789/26952</id>
<updated>2023-12-27T20:07:09Z</updated>
<published>0001-01-01T00:00:00Z</published>
<summary type="text">Regularization in High Dimensional Logistic Regression Model by using Adaptive LASSO Method; การเรกูลาไรซ์ในตัวแบบถดถอยโลจิสติกที่ข้อมูลมีมิติสูงโดยใช้วิธีแลซโซแบบปรับได้
Regularization or penalized logistic regression is widely used to estimate parameters for the high dimensional data. The purpose of this research was to compare the performance of three Adaptive LASSO (Least absolute shrinkage and selection operator)-types for logistic regression in high-dimensional sparse data; Adaptive LASSO using ridge initial weights, and Adaptive LASSO using LASSO initial weights and Adaptive LASSO using Stein-Ridge initial weight and also compared with LASSO under various conditions. The simulation study parameter setting was two cases of sample sizes as n=100, 200, number of quantitative predictors were 2n, 3n, and 4n and additioning 4 binary variables. There are two cases of the relationship structures between predictors and numbers of non-zero regression coefficients of quantitative predictors were 5, 10, and 15 predictors. For each condition, data was iteratively simulated 500 times. For the performance comparison, accuracy of prediction was measured by sensitivity, specificity, and area under ROC curve. The accuracy of parameter estimation was measured by of mean squared error of logistic regression coefficients estimate and variable selection performance was computed by the percentage of non-effected variables including in the model and percentage of effected predictors excluded from the model. The results showed that the Adaptive LASSO method using ridge initial weight had the best performances for all criterion when there are 5 nonzero coefficients of quantitative predictors in the sparse model and the Adaptive LASSO method using Stein-Ridge initial weight had the best performances when there are 10 and 15 nonzero coefficients of quantitative predictors in the sparse model in all cases of other parameters.; การเรกูลาไรซ์หรือการวิเคราะห์การถดถอยโลจิสติกแบบพีนอลไลซ์เป็นวิธีที่ใช้กันอย่างแพร่หลายในการประมาณค่าพารามิเตอร์เมื่อข้อมูลมีมิติสูง งานวิจัยนี้มีวัตถุประสงค์เพื่อทำการศึกษาและเปรียบเทียบประสิทธิภาพการประมาณค่าสัมประสิทธิ์การถดถอยและการคัดเลือกตัวแปรของการเรกูลาไรซ์โดยใช้วิธีแลซโซแบบปรับได้ของตัวแบบถดถอยโลจิสติกในกรณีที่ข้อมูลมีมิติสูงแบบบางเบาทั้ง 3 วิธี ได้แก่ วิธีแลซโซแบบปรับได้ที่ถ่วงน้ำหนักด้วยตัวประมาณริดจ์ (Adaptive LASSO using ridge initial weight), วิธีแลซโซแบบปรับได้ที่ถ่วงน้ำหนักด้วยตัวประมาณแลซโซ (Adaptive LASSO using LASSO initial weight), และวิธีแลซโซแบบปรับได้ที่ถ่วงน้ำหนักด้วยตัวประมาณสไตน์-ริดจ์ (Adaptive LASSO using Stein-Ridge initial weight) รวมทั้งเปรียบเทียบกับวิธีแลซโซ (LASSO) ภายใต้สถานการณ์ที่มีขนาดตัวอย่างคือ 100 และ 200 มีจำนวนตัวแปรอธิบายแบบต่อเนื่องเป็นจำนวน 2, 3, และ 4 เท่าของขนาดตัวอย่าง รวมถึงจำนวนตัวแปรอธิบายแบบไม่ต่อเนื่องจำนวน 4 ตัวแปร รูปแบบความสัมพันธ์ระหว่างตัวแปรอธิบายแตกต่างกัน และจำนวนตัวแปรแบบต่อเนื่องที่สัมประสิทธิ์การถดถอยที่ไม่เท่ากับศูนย์ เท่ากับ 5, 10, และ 15 ตัว ในการจำลองแต่ละสถานการณ์ ทำซ้ำจำนวน 500 รอบ เกณฑ์ที่ใช้ในการทดสอบประสิทธิภาพ คือ ความถูกต้องของการทำนายจากค่าเฉลี่ยของค่าความไว ค่าความจำเพาะ และค่าพื้นที่ใต้โค้งของกราฟ Receiver Operating Characteristic (ROC) ความถูกต้องของการประมาณค่าจากค่าคลาดเคลื่อนกำลังสองเฉลี่ยของค่าประมาณสัมประสิทธิ์การถดถอย และความถูกต้องของการคัดเลือกตัวแปร ผลการวิจัยพบว่า วิธีที่มีประสิทธิภาพดีที่สุดทั้งการทำนาย การประมาณค่า และการคัดเลือกตัวแปร คือ วิธีแลซโซแบบปรับได้ที่ถ่วงน้ำหนักด้วยตัวประมาณวิธีริดจ์ เมื่อมีตัวแปรอธิบายแบบต่อเนื่องที่สัมประสิทธิ์การถดถอยไม่เท่ากับ 0 จำนวน 5 ตัวในตัวแบบบางเบา และวิธีแลซโซแบบปรับได้ที่ถ่วงน้ำหนักด้วยตัวประมาณวิธีสไตน์-ริดจ์ เมื่อมีตัวแปรอธิบายแบบต่อเนื่องที่สัมประสิทธิ์การถดถอยไม่เท่ากับ 0 จำนวน 10 และ 15 ตัวอยู่ในตัวแบบบางเบา ในทุกกรณีของค่าพารามิเตอร์อื่น ๆ
</summary>
<dc:date>0001-01-01T00:00:00Z</dc:date>
</entry>
<entry>
<title>Efficient Comparison of Inference Methods in Logistic Regression Model</title>
<link href="https://sure.su.ac.th/xmlui/handle/123456789/26430" rel="alternate"/>
<author>
<name/>
</author>
<id>https://sure.su.ac.th/xmlui/handle/123456789/26430</id>
<updated>2023-12-05T20:10:59Z</updated>
<published>0012-01-01T00:00:00Z</published>
<summary type="text">Efficient Comparison of Inference Methods in Logistic Regression Model; การเปรียบเทียบประสิทธิภาพของวิธีการอนุมานในตัวแบบถดถอยลอจิสติก
Logistic regression model is used to explain the relationship between explanatory variables and categorical response variable in many research fields. The maximum likelihood (ML) was generally used to estimate parameters, but the ML shows very poor results in the case of separate data. Exact Logistic Regression (ELR) and Markov chain monte carlo (MCMC) exact inference  are the logical alternative to the ML. This research offers a comparison of the property of three methods for estimation. The study found that;

In case of one continuous explanatory variable, when the percentage of the overlapping  ​​is less than or equal to 4, the three estimation methods have low efficient in parameter estimation. When the percentage of overlapping  ​​is higher than 4, the ML and  ELR are equally efficient but the MCMC has the lowest efficiency.                 

In case of two discrete explanatory variables, when percentage of responses is less than 50, the ML and ELR poorly perform in parameter estimation. However, when the percentage of response is equal to 50 at sample size less than or equal to 28, the ELR is more effective than the ML, but at sample size greater than or equal to 48, both methods are similarly effective. 

In interval estimation of two cases the results showed that, when the probability estimation coverage does not differ from the confidence level, the least confidence interval width estimation is ML. The ELR and MCMC methods commondly provided the infinite width of the confidence interval on average.; ตัวแบบการถดถอยลอจิสติกได้นำมาใช้อย่างกว้างขวางในงานวิจัยหลายแขนง เพื่ออธิบายความเกี่ยวพันระหว่างตัวแปรอธิบายและตัวแปรตอบสนอง ในกรณีที่ตัวแปรตอบสนองเป็นตัวแปรจำแนกประเภท โดยทั่วไปวิธีการประมาณค่าพารามิเตอร์ในตัวแบบการถดถอยลอจิสติกที่นิยมใช้ คือ วิธีภาวะน่าจะเป็นสูงสุด (ML) แต่จากการศึกษาพบว่าวิธีดังกล่าวไม่สามารถใช้ในการประมาณค่าพารามิเตอร์ในกรณีข้อมูลเป็นแบบแบ่งแยกสมบูรณ์ และข้อมูลแบบแบ่งแยกกึ่งสมบูรณ์ งานวิจัยนี้ศึกษาวิธีการประมาณค่าพารามิเตอร์วิธีอื่นที่สามารถแก้ไขข้อบกพร่องของวิธีภาวะน่าจะเป็นสูงสุด ได้แก่ วิธีการถดถอยลอจิสติกแบบแม่นตรง (ELR) และวิธีลูกโซ่มาร์คอฟมอนติคาร์โล (MCMC) สำหรับการถดถอยลอจิสติกแบบแม่นตรง จากนั้นทำการเปรียบเทียบประสิทธิภาพของวิธีการประมาณค่าพารามิเตอร์ทั้งสามวิธี เมื่อข้อมูลมีจำนวนค่าสังเกตทับซ้อนต่างกัน และกรณีที่ข้อมูลมีร้อยละของค่าตอบสนองต่างกัน จากการศึกษาพบว่า

กรณีตัวแปรอธิบายแบบต่อเนื่อง 1 ตัวแปร: สำหรับทุกร้อยละของค่าตอบสนอง เมื่อร้อยละของจำนวนค่าสังเกตทับซ้อนมีค่าต่ำกว่าหรือเท่ากับ 4  วิธีการประมาณค่าพารามิเตอร์ทั้ง 3 วิธี ยังมีประสิทธิภาพในการประมาณค่าพารามิเตอร์ต่ำ เมื่อร้อยละของจำนวนค่าสังเกตทับซ้อนมีค่าสูงกว่า 4 วิธี ML และวิธี ELR มีประสิทธิภาพใกล้เคียงกัน ส่วนวิธี MCMC มีประสิทธิภาพต่ำสุด

กรณีตัวแปรอธิบายแบบไม่ต่อเนื่องทั้ง 2 ตัวแปร: เมื่อค่าประมาณร้อยละของค่าตอบสนองต่ำกว่า 50 วิธี ML และ ELR ยังมีประสิทธิภาพต่ำในการประมาณค่าพารามิเตอร์ เมื่อค่าประมาณร้อยละของค่าตอบสนองเท่ากับ 50 ที่ขนาดตัวอย่างน้อยกว่าหรือเท่ากับ 28 วิธี ELR มีประสิทธิภาพสูงกว่าวิธี ML ส่วนที่ขนาดตัวอย่างมากกว่าหรือเท่ากับ 48 ทั้งสองวิธีมีประสิทธิภาพใกล้เคียงกัน

การประมาณค่าพารามิเตอร์แบบช่วงของทั้ง 2 กรณี พบว่า วิธีที่ให้ค่าประมาณค่าความน่าจะเป็นครอบคลุมไม่แตกต่างจากระดับความเชื่อมั่นที่กำหนดและให้ค่าประมาณความกว้างเฉลี่ยของช่วงความเชื่อมั่นน้อยที่สุด คือวิธี ML เนื่องจากวิธี ELR และวิธี MCMC โดยส่วนใหญ่ให้ค่าประมาณความกว้างเฉลี่ยของช่วงความเชื่อมั่นเป็นอนันต์
</summary>
<dc:date>0012-01-01T00:00:00Z</dc:date>
</entry>
</feed>
