Hiển thị các bài đăng có nhãn statistics. Hiển thị tất cả bài đăng
Hiển thị các bài đăng có nhãn statistics. Hiển thị tất cả bài đăng

Thứ Tư, 14 tháng 7, 2010

Uyển ngữ, số thống kê, và chất vấn


Lại là một trong những cái tựa linh tinh trên blog này. Nhưng nó phản ánh cái "muôn mặt", hoặc "nhốn nháo", của cuộc sống quanh tôi.

Trước hết, hãy nói về uyển ngữ. Uyển ngữ, là từ tiếng Việt dùng để dịch từ euphemism trong tiếng Anh. Chẳng hiểu tôi đã học được từ này (tiếng Việt) ở đâu, nhưng chắc chắn là học sau năm 1975, có lẽ vào khoảng cuối thập niên 80 thì phải. Lúc ấy, tôi còn đang dạy ở Khoa Ngữ Văn Anh trường ĐHKHXH-NV.

Tôi vẫn nhớ cảm giác của lần đầu tiên khi tôi gặp được từ này. Rất thích thú! Đọc lên là hiểu ngay lập tức. Vì nó ngắn gọn, và ... rất "uyển ngữ"! Tôi nhớ, chẳng cần tra từ điển mà tôi đã suy ngay ra được nghĩa của từ này, do suy ra từ các từ có chứa gốc "uyển" mà tôi biết (uyển chuyển, vườn thượng uyển).

Nếu không có từ này, thực sự không biết làm sao để dịch từ euphemism trong tiếng Anh sang tiếng Việt cho gọn ghẽ nhỉ?

Dưới đây là định nghĩa của từ euphemism, lấy trong OALD (2000):

an indirect word or phrase that people often use to refer to sth embarrassing or unpleasant, sometimes to make it seem more acceptable than it really is; ex: "pass away" is a euphemism for "die"; "user fees" is just a politician's euphemism for taxes.

Rõ rồi nhé: euphemism là nói nhẹ đi, nói tránh, để cho dễ chấp nhận. Chẳng hạn, "chết" thì nói là "qua đời". Trước đây, hồi còn trẻ, đang học đại học, tôi và một nhóm bạn hay đến thăm trường mù Nguyễn Đình Chiểu, tạo được những mối quan hệ thân thiết với các em ở đây.

Lần đầu tiên đến đấy, tôi được dặn không được nói "mù", mà phải nói "khiếm thị" (trịnh trọng), hoặc bình dân hơn, là "không thấy đường", ví dụ như nói: Tôi có quen một anh bạn kia, anh ấy "không thấy đường", nhưng tốt bụng lắm! Đấy cũng là uyển ngữ.

Vậy uyển ngữ chính là cách sử dụng ngôn ngữ một cách khéo léo, để tạo ra hiệu quả mong muốn. Rõ ràng là tốt. Và mọi người đều phải học cách dùng nó, nếu muốn là một người sử dụng ngôn ngữ thành công.

Vậy thì có gì để mà nói chứ? Nghiên cứu về uyển ngữ là việc của một nhà văn hóa, hoặc một nhà ngôn ngữ, chứ không phải là việc của tôi, lại càng không phải là "nhiệm vụ chính trị" của cái blog này, vốn là "cõi riêng" để tôi ghi chép linh tinh cho riêng mình và bạn bè, thân hữu. Tôi không viết cái entry này để hô hào mọi người phải học và dùng uyển ngữ chứ? Ấy chết, hoàn toàn không phải thế!

Các bạn còn nhớ là cách đây ít lâu tôi đã viết về "thổ tả" phải không? Thế các bạn có nhớ rằng hôm ấy tôi nhắc đến từ "euphemism" chăng? Nó là lần đầu tôi nhắc đến uyển ngữ trên blog này đấy. Nhắc đến, và ... rất cáu!!!!! Rồi tôi quên đi.

Nhưng cách đây vài hôm, tôi lại đọc được trên blog của Huy Quang Piano (huyquangpiano.blogspot.com), một người Hà Nội, cách nói của các quan chức Hà Nội, cũng là ví dụ tiêu biểu của "thuật dùng uyển ngữ". Họ nói cái gì thế?

Xin đọc đoạn dưới đây:
Mưa không to lắm nhưng đường phố biến thành sông.
Tất nhiên là hệ thống thoát nước thành phố có vấn đề, nhưng khổ nỗi, người ta lại bảo "cơ sở hạ tầng không phát triển kịp với tốc độ đô thị hoá" (!).

Phần in đậm đậm mà tôi thêm vào chính là cái uyển ngữ của quan chức Hà Nội đấy. Quá xuất sắc phải không, đúng là bậc thầy trong việc dùng uyển ngữ.

Chỉ có điều, uyển ngữ đâu có phải là việc của các nhà quản lý, nhất là khi nói chuyện với con dân, những người dốt nát, chỉ có thể hiểu khi mọi việc được nói toạc ra thẳng thắn. Tôi đang nói về chúng ta ấy, những người đọc blog này, phải vậy không các bạn?

Sẽ có người hỏi tôi, nhưng nếu có nhiều vấn đề dân chúng bức xúc, mà quan chức không được dùng uyển ngữ, vậy thì dùng cái gì? À, thì dùng số thống kê, chứ sao! Con người "quản lý" của tôi trả lời thế. Chứ còn gì nữa? Cho nó rõ ràng, chính xác, minh bạch chớ!

Rồi tôi bỗng nhớ đến số thống kê của VN. Và nhớ rằng, chính mình đã viết một entry có tên là "Số thống kê, trời ơi!" trên blog này. Nói về những thống kê về tỷ lệ tốt nghiệp THPT trong năm nay.

Gần đây, do có nhiều số thống kê về giáo dục khá là ... khó hiểu, tôi phải tìm mua qua amazon và đang đọc một cuốn sách có tựa là "How to lie with statistics".

Cuốn sách này mỏng, nhỏ, và nội dung thường thôi, nhưng cuối cuốn sách có một chương rất khá, có đưa ra lời khuyên với bạn đọc khi đọc số thống kê do người khác đưa ra. Lời khuyên ấy là: khi đọc một số thống kê thì đừng tin ngay, mà phải tự hỏi mình 5 câu hỏi, mà tôi tự tóm tắt cho dễ nhớ như sau:

1. Ai nói? Có thiên vị không? 2. Sao họ biết mà nói? 3. Số liệu có đủ không? 4. Số liệu có ăn nhập với kết luận không? 5. Kết luận có lý không?

Đọc xong thì thấy ngay câu trả lời hiển nhiên với những con số thống kê về kỳ thi tốt nghiệp có đáng tin hay không. Đây nhé:

1. Người nói là Bộ Giáo dục, tất nhiên sẽ có khả năng thiên vị;

2. Người nói không phải là người trực tiếp làm, nên nếu muốn biết sự thật thì chỉ có các trường, thầy cô, phụ huynh và chính học sinh mới biết thật ra chất lượng có tăng không mà thôi;

3. Số liệu không đủ, vì chỉ có kết quả thi, không nói gì vể độ khó của đề thi, về cách định cỡ bài thi để đảm bảo tương đương giữa 2 kỳ thi, vv;

4. Số liệu không ăn nhập với kết luận, vì điểm của một kỳ thi thì chỉ nói lên kỳ thi đó khó hay dễ đối với thí sinh của năm đó, còn chất lượng giáo dục là chất lượng giáo dục, hai việc không phải và không thể là một;

5. Kết luận về chất lượng GD đã tăng chẳng có lý tý nào, vì nếu cứ tăng mãi với tốc độ này thì sang năm sẽ thế nào? Tỷ lệ đậu là 120% chắc?????

Thế mà người ta vẫn nói được đấy! Lại còn dùng những số thống kê ấy để chất vấn người khác nữa chứ!

À, chất vấn!!! Mới có một vụ chất vấn nổi đình đám kia kìa, kèm với một vụ từ chức (à, thì chưa làm đơn, nhưng cũng đã buột miệng nói ra, vì mệt mỏi quá!) Ai không biết vụ đình đám đó là gì, cứ hỏi google với những từ "Giám đốc Sở Giáo dục tỉnh Bình Dương từ chức", sẽ rõ!

Người ta trách ông ... nóng nảy! Ừ, mà nóng nảy thật! Chất vấn là việc của Hội đồng, ông làm chính quyền thì ông phải chịu chất vấn chứ! Cứ từ từ mà trả lời, rồi mọi người sẽ hiểu và ủng hộ. Chứ ông không thấy còn bao nhiêu việc khác, cái gì mà 5 cái cổng chào tốn bao nhiêu tỷ gì đó ở Hà Nội, người ta có từ chức đâu? Có gì mà nóng quá vậy?

Nhưng mà nóng thiệt! Vì người ta chất vấn ông bằng chính những cái thống kê ... trời ơi đất hỡi trên kia kìa! Gặp tôi, thì tôi cũng từ chức. Chịu gì nổi!

Làm gì phải từ chức! Tôi nghe bạn bè tôi nói thế. Phải dấn thân, phải dũng cảm chứ!

Vậy nếu làm đàng hoàng, mà chẳng ai ủng hộ, bị chất vấn bằng mấy con số thống kê ... thổ tả như GĐ Sở GD Bình Dương kia, thì phải làm sao? Tôi tự hỏi.

Rồi chợt ngộ ra câu trả lời: Dùng uyển ngữ, chứ còn gì nữa! Chỉ có thế thôi, mà nghĩ mãi không ra! Vậy mà ở trên tôi dám nói là không định viết entry này để hô hào dùng uyển ngữ! Đúng là sai lầm!

Vậy xin được nói lại: Mọi người đều cần cách dùng uyển ngữ, và dùng số thống kê để ... nói dối. Cần mượn sách How to lie with statistics thì liên hệ với tôi.

Mà chắc cũng chẳng ai cần, có khi VN còn dạy được thế giới làm điều này nữa ấy chứ!

Tự nhiên tôi nhớ Gabriel Marcia Marquez quá đỗi!
--
Cập nhật lúc 10:17 phút cùng ngày:
Hai cái hình này là hình về mưa, lụt tại TQ, lấy trên xinhuanet.com. Theo yêu cầu của Khuê, vì "viết thì phải có hình minh họa, cho nó sinh động chứ mẹ!"

Ờ mà uyển ngữ, có phải là từ Hán Việt không? Vậy biệt tài dùng uyển ngữ, chắc phải có nguồn gốc từ Trung Quốc nhỉ?

Thứ Bảy, 30 tháng 5, 2009

statistics jokes

In God we trust. All others must bring data.

If had only one day left to live, I would live it in my statistics class:
it would seem so much longer.

The latest survey shows that 3 out of 4 people make up 75% of the world's population.

Three statisticians went out hunting, and came across a large deer. The first statistician fired, but missed, by a meter to the left. The second statistician fired, but also missed, by a meter to the right. The third statistician didn't fire, but shouted in triumph, "On the average we got it!"

Did you hear about the politician who promised that, if he was elected, he'd make certain that everybody would get an above average income?

There are three kinds of lies: lies, damned lies, and statistics.

Logic is a systematic method for getting the wrong conclusion with confidence.
Statistics is a systematic method for getting the wrong conclusion with 95% confidence.

A statistics major was completely hung over the day of his final exam. It was a True/False test, so he decided to flip a coin for the answers. The stats professor watched the student the entire two hours as he was flipping the coin...writing the answer...flipping the coin...writing the answer. At the end of the two hours, everyone else had left the final except for the one student. The professor walks up to his desk and interrupts the student, saying:
"Listen, I have seen that you did not study for this statistics test, you didn't even open the exam. If you are just flipping a coin for your answer, what is taking you so long?"
The student replies bitterly, as he is still flipping the coin: "Shhh! I am checking my answers!"

There was this statistics student who, when driving his car, would always accelerate hard before coming to any junction, whizz straight over it , then slow down again once he'd got over it. One day, he took a passenger, who was understandably unnerved by his driving style, and asked him why he went so fast over junctions. The statistics student replied, "Well, statistically speaking, you are far more likely to have an accident at a junction, so I just make sure that I spend less time there."

Three professors (a physicist, a chemist, and a statistician) are called in to see their dean. Just as they arrive the dean is called out of his office, leaving the three professors there. The professors see with alarm that there is a fire in the wastebasket.
The physicist says, "I know what to do! We must cool down the materials until their temperature is lower than the ignition temperature and then the fire will go out."
The chemist says, "No! No! I know what to do! We must cut off the supply of oxygen so that the fire will go out due to lack of one of the reactants."
While the physicist and chemist debate what course to take, they both are alarmed to see the statistician running around the room starting other fires. They both scream, "What are you doing?"
To which the statistician replies, "Trying to get an adequate sample size."

I asked a statistician for her phone number... and she gave me an estimate.

A statistician's wife had twins. He was delighted. He rang the minister who was also delighted.
"Bring them to church on Sunday and we'll baptize them," said the minister.
"No," replied the statistician. "Baptize one. We'll keep the other as a control.

Thứ Bảy, 21 tháng 3, 2009

Probability and the law

Problem Solving – a Statistician’s Guide – 2nd ed by Chris Chatfield
Chapman & Hall 1988, 1995

Exercise G1 – Probability and the law
(p. 215)
In a celebrated criminal case in California (People versus Collins, 1968), a black male and a white female were found guilty of robbery, partly on the basis of a probability argument. Eyewitnesses testified that the robbery had been committed by a couple consisting of a black man with a beard and a moustache, and a white woman with blond hair in a ponytail. They were seen driving a car which was partly yellow. A couple, who matches these descriptions, were later arrested. In court they denied the offence and could not otherwise be positively identified.

A mathematics lecturer gave evidence that the six main characteristics had probabilities as follows:

Negro man with beard 1/10
Man with moustache 1/4
Girl with pony tail 1/10
Girl with blond hair 1/3
Partly yellow car 1/10
Inter-racial couple in car 1/1000

The witness then testified that the product rule of probability could be used to multiply these probabilities together to give a probability of 1/12 000 000 that a couple chosen at random would have all these characteristics. The prosecutor asked the jury to infer that there is only one chance in 12 million of the defendants’ innocence, and the couple were subsequently convicted.

Comment on the above probability argument.


(Please come back next week for the solution!)

Thứ Hai, 16 tháng 3, 2009

Đo lường và lãng phí thông tin - Nguồn: Lao động cuối tuần 15/3/09

(LĐCT) - Khủng hoảng kinh tế làm cho nhiều người bị thất nghiệp. Số người bị mất việc làm trong một khoảng thời gian nào đó, thí dụ trong hai tháng đầu hay trong quý một của năm 2009, là con số rất quan trọng làm cơ sở để hiệu chỉnh các chính sách của Nhà nước và cho định hướng hoạt động của các doanh nghiệp.
Đo lường chính xác các con số như vậy không đơn giản nếu không nói là không thể. Tuy nhiên, có các phương pháp thống kê để có thể ước lượng những con số như vậy. Xây dựng các hệ thống đo lường như thế một cách hiệu quả là việc rất quan trọng để tăng hiệu quả quản trị đất nước.

Tại các nước phát triển, người ta đã xây dựng được các hệ thống như thế và định kỳ họ công bố công khai các số liệu đo lường hay ước lượng như vậy. Đấy là một phần quan trọng trong cơ sở hạ tầng thông tin có vai trò hết sức to lớn trong sự phát triển của mọi nền kinh tế.

Thí dụ, theo Cục Thống kê Lao động thuộc Bộ Lao động Mỹ, số chỗ làm việc bị mất (ngoài ngành nông nghiệp) của nước này trong tháng 2.2009 là 651 ngàn chỗ và làm cho số người thất nghiệp tăng từ 11 triệu 616 ngàn người (tháng 1.2009) lên 12 triệu 467 ngàn người (tháng 2.2009). Cơ quan này cũng công bố công khai số liệu từng tháng của nhiều năm. Các nước phát triển khác cũng vậy. Và điều quan trọng là các số liệu này rất dễ tiếp cận trên mạng.

Rất tiếc tại Việt Nam chưa có hệ thống đo lường như vậy. Bộ Lao động Thương binh và Xã hội không có cục thống kê, tuy có trang www.vieclamvietnam.gov.vn nhưng thực ra chỉ là trang dịch vụ việc làm có chất lượng chưa tương xứng; Tổng cục Thống kê có Vụ Thống kê Dân số và Lao động nhưng chức năng khá chung chung.

Không rõ vụ này có tiến hành điều tra, đo lường và công bố kết quả thường xuyên không. Nhiều khả năng là không vì rất khó tìm các số liệu chính thức như vậy. Còn các số liệu "dự báo" về thất nghiệp luôn đá nhau: người dự báo số thất nghiệp tăng lên trong năm 2009 là vài triệu, Bộ LĐ-TB và XH cho là khoảng 150-300 ngàn (tính theo mức giảm "chỉ tiêu" tăng trưởng GDP để suy ra số việc làm mất đi), người thì ước tính nửa triệu người...

Các số liệu "dự báo" như thế khó có thể lấy làm cơ sở cho hiệu chỉnh chính sách. Dự báo là một chuyện (về tương lai), đo lường cái thực sự đã xảy ra (trong quá khứ) là chuyện khác. Việc hoạch định chính sách nên dựa cả vào số liệu đo lường và dự báo. Việc xây dựng các hệ thống đo lường như vậy và việc công bố công khai các số liệu là một đòi hỏi rất bức bách của sự phát triển đất nước.

Trên đây mới chỉ nói đến số liệu về thất nghiệp. Còn nhiều số liệu khác mà chúng ta cần đo lường và công bố. Thực ra, từ đổi mới đến nay Tổng cục Thống kê đã có những nỗ lực vượt bậc và rất đáng trân trọng trong việc đo lường và công bố các số liệu thống kê về kinh tế và xã hội song so với các nước khác trên thế giới và trong khu vực thì còn quá sơ sài và ít ỏi.

Nước ta đang trong quá trình hội nhập và muốn hội nhập nhanh và hiệu quả thì rất nên dùng các khái niệm mà cả thế giới đều dùng và nên tránh "sáng tạo" ra những khái niệm "mới" chẳng giống ai, hay sử dụng các khái niệm thông dụng với ý nghĩa khác. Vì những việc như vậy sẽ gây hiểu nhầm, cản trở sự giao lưu và cản trở sự phát triển.

Một thí dụ điển hình là khái niệm "xã hội hóa" đã được bàn tới nhiều. Liên quan đến số liệu thống kê cách dùng "vốn đầu tư" cũng vậy. Vốn và đầu tư là hai khái niệm hoàn toàn khác nhau. Việc trộn lẫn hai thứ trong cách dùng "vốn đầu tư" gây nhiều nhầm lẫn và rắc rối. Cách tính hệ số ICOR (hệ số cho biết cần bao nhiêu đồng đầu tư để tạo ra thêm một đồng GDP) của chúng ta cũng vậy.

Một số người đề ra 3 cách tính ICOR khác nhau và cho những kết quả rất khác nhau. Mỗi người tính một phách, gây rắc rối trong đánh giá hiệu quả đầu tư. Các cơ quan quốc tế như Ngân hàng Thế giới, Quỹ Tiền tệ Quốc tế, UNDP cũng đã khuyên chúng ta dùng "ngôn ngữ thống kê" chung như vậy.

Đáng tiếc vẫn còn nhiều thứ chúng ta dùng khác với thế giới và điều đó gây khó khăn cho sự so sánh quốc tế, cho việc đánh giá thành tích phát triển của chính chúng ta, cho công việc của các nhà nghiên cứu. Đấy là một sự lãng phí lớn về thông tin.

Lại có loại thông tin chúng ta đã bỏ khá nhiều tiền và công sức ra thu thập và đo lường, song việc công bố chúng không đầy đủ hay chậm trễ làm cho việc sử dụng chúng kém hiệu quả. Tổng cục Thống kê thường xuyên tiến hành điều tra về mức sống hộ gia đình (số liệu đã qua xử lý của các năm 2002, 2004 và 2006 đã được xuất bản), trong phần lao động việc làm không sao tìm thấy số liệu về thất nghiệp.

Những số liệu về thu nhập và chi tiêu liệu có thể sử dụng để tính toán hay làm cơ sở cho các chính sách kích cầu? Tôi e là khó và nếu có thể thì có lẽ cũng đã chẳng ai dùng. Số liệu công bố là số liệu đã được xử lý, liệu có cơ chế nào để tiếp cận đến số liệu nguồn (chưa xử lý) để có thể chắt lọc ra những thông tin xác đáng làm cơ sở cho hoạch định chính sách? Những thông tin điều tra về doanh nghiệp cũng vậy.

Đấy là những thông tin rất có giá trị để hoạch định chính sách phát triển kinh tế, để cho các doanh nghiệp có thể định hướng phát triển. Sử dụng kho báu này ra sao là việc hệ trọng, nếu không được dùng một cách thích hợp chúng chỉ là dữ liệu "chết". Nếu có cơ chế thông thoáng để tiếp cận đến những thông tin đã được thu thập và đo lường này thì những dữ liệu ấy không chỉ là số liệu "chết" mà thực sự là nguồn tài nguyên rất quý giá cần được khai thác hiệu quả cho sự phát triển của đất nước.

"Cát cứ thông tin", dùng thông tin mà mình được uỷ thác cai quản để trục lợi (như ém nhẹm thông tin quy hoạch và cấp cho các nhà đầu cơ nhà đất làm giàu nhanh chóng qua chạy dự án hay "buôn" chính sách là những hiện tượng được nhiều người nhắc tới) đang là một vấn nạn cần xoá bỏ.

Các hệ thống đo lường thông tin kinh tế-xã hội, việc công bố thông tin, cơ chế tiếp cận thông tin, sử dụng hữu hiệu và tránh phung phí tài nguyên thông tin là những nhân tố quan trọng cho sự phát triển đất nước, đã được chương trình quốc gia về công nghệ thông tin đề xuất từ cả chục năm nay nhưng chưa được chú ý đúng mức. Cần khẩn trương xây dựng hạ tầng thông tin bao gồm các hệ thống đo lường, xử lý, cung cấp thông tin như vậy.

Nguyễn Quang A

Thứ Bảy, 14 tháng 3, 2009

Bài tập Thống kê ứng dụng

(trích trong chương 1 Seeing through statistics)
*8. Suppose you have a choice of two grocery stores in your neighborhood. Because you hate waiting, you want to choose the one for which there is generally a shorter wa cit in the check in line. How would you gather information to determine which one is faster? Would it be sufficient to visit each store once and time how long you had to wait in line? Explain.
*8. Giả sử ở khu vực bạn ở có hai tiệm tạp hóa. Vì không thích chờ đợi nên bạn muốn chọn tiệm nào ít phải chờ hơn. Làm sao bạn có thể thu thập thông tin thể xác định nơi nào tốt hơn? Đến mỗi tiệm một lần và tính thử xem phải chờ bao lâu đã đủ chưa? Hãy giải thích.

*9. Suppose researchers want to know whether smoking cigars increases the risks of esophageal cancer.
a. Could they conduct a randomized experiment to test this? Explain.
b. If they conducted an observational study and found that cigar smokers had a higher rate of esophageal cancer than those who did not smoke cigars, could they conclude that smoking cigars increases the risk of esophageal cancer? Explain why or why not?

*9. Giả sử các nhà nghiên cứu muốn tìm hiểu xem việc hút thuốc có làm tăng nguy cơ bị bệnh ung thư cuống họng không.
a. Có thể thực hiện một cuộc thử nghiệm theo xác xuất để kiểm tra điều này không? Tại sao?
b. Nếu tiến hành một nghiên cứu quan sát và thấy rằng những người hút thuốc có tỷ lệ ung thư cao hơn những người không hút thuốc, có thể kết luận rằng hút thuốc làm tăng nguy cơ ung thư cuống họng không? Hãy giải thích.


*13. Suppose you have 20 tomato plants and want to know if fertilizing them will help them produce more fruit. You randomly assign 10 of them to receive fertilizer and the remaining 10 to receive none. You otherwise treat the plants in an identical manner.
a. Explain whether this would be observational study or a randomized experiment.
b. If the fertilized plants produce 30% more fruit than the unfertilized plants, can you conclude that fertilized caused the plants to produce more? Explain.

*13. Giả sử bạn có 20 cây cà chua và cần biết nếu bón phân thì có làm tăng năng suất không. Bạn phân nhóm ngẫu nhiên (xác suất) 10 cây có bón phân và 10 cây không bón phân. Cả 20 cây này đều được đối xử như nhau trên các bình diện khác.
a. Hãy giải thích đây là nghiên cứu quan sát hay nghiên cứu xác suất?
b. Nếu cây đượ c bón phân có năng suất cao hơn 30%, có thê kết luận là điều này do bón phân không? Tại sao?

*15. National polls are often conducted by asking the opinions of a few thousand adults nationwide and using them to infer the opinions of all adults in the nation. Explain who is in the sample and who is in the population for such polls.

*15. Các cuộc điều tra cấp quốc gia thường được thực hiện bầng cách hỏi ý kiến vài ngàn người trên cả nước. Như vậy ở đây ai là mẫu và ai là dân số.

*17. Suppose a study first asked people whether they meditate regularly and then measured their blood pressures. The idea would be to see if those who meditate have lower blood pressure than those who do not do so.
a. Explain whether this would be an observational study or a randomized experiment.
b. If it were found that meditators had lower-than-average blood pressures, can we conclude that meditation causes lower blood pressure? Explain

Learning outcomes (chuẩn đầu ra) của môn học Thống kê ứng dụng nhập môn

Môn Thống kê ứng dụng nhập môn nhằm trang bị cho người học các kỹ năng sau:
1. tóm tắt và trình bày số liệu sẵn có sao cho thể hiện được tính quy luật và bản chất của hiện tượng được quan sát --> thống kê mô tả
2. thu thập hoặc tạo ra số liệu mới về một hiện tượng cần quan sát để bộc lộ bản chất của hiện tượng --> thiết kế các nghiên cứu định lượng (kỹ năng này sẽ được củng cố về mặt lý luận trong môn học Nghiên cứu khoa học giáo dục)
3. diễn giải ý nghĩa số liệu dựa trên mẫu nhỏ để rút ra quy luật của tổng thể --> thống kê suy diễn

Thứ Bảy, 28 tháng 2, 2009

Thống kê và nghiên cứu khoa học

NGHIÊN CỨU KHOA HỌC LÀ GÌ?
Là tìm lời giải đáp cho những câu hỏi/ vấn đề chưa có câu trả lời bằng cách thu thập, phân tích diễn giải các chứng cứ khoa học theo các nguyên tắc và phương pháp khoa học phù hợp.

Quotes:
Research is the cornerstone of any science, including both the hard sciences such as chemistry or physics and the social (or soft) sciences such as psychology, management, or education. It refers to the organized, structured, and purposeful attempt to gain knowledge about a suspected relationship.

http://allpsych.com/researchmethods/introduction.html

Research is a process of steps used to collect and analyse information in order to increase our understanding of a topic of issue. At a general level, research consists of three steps:
1. Pose a question
2. Collect data to answer the question
3. Present an answer to the question
This should be a familiar process. You engage in solving problem everyday and you start with a question, collect some information, and then form an answer. Although there are a few more steps than these three, this is the overall framework for reasearch. When you examine an published study or conduct your own study, you will find these three parts as the core elements.
WHAT ARE STEPS IN CONDUCTING RESEARCH?
When researchers conduct a study, they proceed through a distinct set of steps. Years ago these steps were identified as the “scientific method” of inquiry (Kerlinger, 1972; Leedy & Ormrod, 2001). Using a “scientific method”, reseachers:
• Identify a problem that defines the goal of research
• Make a prediction that, if confirmed, resolves the problem
• Gather date relevant to this prediction
• Analyze and interpret the data to see if it supports the prediction and resolves the question that initiated the research.
Applied today, theses steps provide the foundation for educational research. Although not all studies include predictions, you engage in these steps whenever you undertake a research study. As shown in Figure 1.2, the process of research consists of six steps:
1. Identifying a research problem
2. Reviewing the literature
3. Specifying a purpose for research
4. Collecting data
5. Analyzing and interpreting the data

6. Reporting and evaluating research.

(From: pp 3-8, Chapter 1, Educational Research – Planning, Conducting, and Evaluating Quantitative and Qualitative Research, John W. Creswell, 2nd ed, Pearson 2005, 2002)

CÁC PHƯƠNG PHÁP THU THẬP VÀ XỬ LÝ DỮ LIỆU TRONG NGHIÊN CỨU KHOA HỌC

+ phương pháp định lượng: sử dụng các con số
+ phương pháp định tính: sử dụng các chứng cứ không phải là số (ngôn ngữ hình ảnh âm thanh)

Quotes:
Quantitative versus Qualitative Research Methods

There are essentially two forms of educational research: quantitative (or statistical) and qualitative. In the past, most research was quantitative in nature, fashioned after the highly successful hard-science research methods. Since the 1990s, educational researchers have embraced qualitative research with the recognition that research on the human mind is fundamentally different from research on physical systems. Qualitative methods are used for depth of knowledge and quantitative methods are used for breadth and generalizability (emphasis added by PA). We cannot emphasize this enough: the best research designs employ both methods.
Most readers of this article will be familiar and comfortable with the ideas behind quantitative research. Because of this, and because quantitative research methods are well documented (e.g., Hopkins, 1998 ), we will not attempt a discussion of statistical research methods here. However, many readers will be uncomfortable with qualitative research because it is so different from the formal training of a biologist. Nonetheless, it is a vital component of educational research and should not be overlooked. What follows is a brief discussion of the most useful and common qualitative research methods for education.

Qualitative Research

Statistical research is well suited for drawing generalizable and repeatable conclusions based on a large sample of students, but it lacks depth. For example, suppose you use a particular approach in your course that you believe will lead to a greater understanding of transcription. You design a test that measures student understanding and give it to numerous students taught by using your approach and a traditional approach. You find that the students taught under your new method significantly outperform their peers. You come to the conclusion that your approach is superior. If your test and research methods were well designed, your conclusion is valid. You know your method is superior, but you have no data that tell you why the method was so successful (emphasis added by PA).. You probably have hypotheses about why the method worked so well (you had some reason to try it in the first place), but you do not know for sure. To properly answer the question of why, you must engage in qualitative research.

Qualitative research is also enormously helpful when you are initially investigating a particular area and are in the process of determining what questions might be interesting to study. The data of qualitative research are incredibly rich. Although this richness is an asset, it also makes qualitative data difficult and time consuming to analyze.

The National Science Foundation (NSF) has developed a detailed introduction to qualitative research geared toward the scientist. This introduction, User-Friendly Handbook for Mixed Method Evaluations, is excellent and can be found online at www.nsf.gov/pubsys/ods/getpub.cfm?nsf97153.

http://www.pubmedcentral.nih.gov/articlerender.fcgi?artid=128540&tool=pmcentrez

THỐNG KÊ LÀ GÌ
PA: Là các nguyên tắc và phương pháp (quy trình, thủ tục) thu thập và phân tích các dữ liệu số (numerical data) về các hiện tượng để tìm hiểu bản chất và quy luật của các hiện tượng đó.

Quotes:
A statistic is a numerical representation of information. Whenever we quantify or apply numbers to data in order to organize, summarize, or better understand the information, we are using statistical methods. These methods can range from somewhat simple computations such as determining the mean of a distribution to very complex computations such as determining factors or interaction effects within a complex data set.

http://allpsych.com/researchmethods/descriptivestatistics.html

The science of statistics deals with the collection, analysis, interpretation, and presentation of data. We see and use data in our everyday lives. To be able to use data correctly is essential to many professions and in your own best self-interest.
(Collaborative statistics)

Statistics consists of the principles and methods for
1. Designing studies
2. Collecting data
3. Presenting and analysing data
4. Interpreting the results
Statistics has been described as
1. Turning data into information
2. Data-based decision making
3. The technology of the "Scientific Method"

Surfstat.australia: an online text in introductory Statistics
intro1.html 02/21/2009 16:23:37

HỌC XONG THỐNG KÊ CÓ THỂ LÀM ĐƯỢC GÌ (LEARNING OUTCOMES)?
Trả lời dựa trên trang web của surfstats

http://surfstat.anu.edu.au/surfstat-home/surfstat-main.html
- Tóm tắt và trình bày số liệu
- Tạo số liệu mới (thiết kế nghiên cứu định lượng)
- Suy luận thống kê


SO SÁNH ƯU VÀ NHƯỢC ĐIỂM CỦA HAI PHƯƠNG PHÁP ĐỊNH LƯỢNG VÀ ĐỊNH TÍNH

What Are the Major Differences Between Quantitative and Qualitative Techniques?
As shown in Exhibit 1, quantitative and qualitative measures are characterized by different techniques for data collection.
Exhibit 1. Common techniques
Quantitative Qualitative
Questionnaires
Tests
Existing databases Observations
Interviews
Focus groups


Aside from the most obvious distinction between numbers and words, the conventional wisdom among evaluators is that qualitative and quantitative methods have different strengths, weaknesses, and requirements that will affect evaluators’ decisions about which methodologies are best suited for their purposes. The issues to be considered can be classified as being primarily theoretical or practical.
Theoretical issues. Most often, these center on one of three topics:
• The value of the types of data;
• The relative scientific rigor of the data; or
• Basic, underlying philosophies of evaluation.

Value of the data. Quantitative and qualitative techniques provide a tradeoff between breadth and depth and between generalizability and targeting to specific (sometimes very limited) populations. For example, a sample survey of high school students who participated in a special science enrichment program (a quantitative technique) can yield representative and broadly generalizable information about the proportion of participants who plan to major in science when they get to college and how this proportion differs by gender. But at best, the survey can elicit only a few, often superficial reasons for this gender difference. On the other hand, separate focus groups (a qualitative technique) conducted with small groups of male and female students will provide many more clues about gender differences in the choice of science majors and the extent to which the special science program changed or reinforced attitudes. But this technique may be limited in the extent to which findings apply beyond the specific individuals included in the focus groups.
Scientific rigor. Data collected through quantitative methods are often believed to yield more objective and accurate information because they were collected using standardized methods, can be replicated, and, unlike qualitative data, can be analyzed using sophisticated statistical techniques. In line with these arguments, traditional wisdom has held that qualitative methods are most suitable for formative evaluations, whereas summative evaluations require "hard" (quantitative) measures to judge the ultimate value of the project.

This distinction is too simplistic. Both approaches may or may not satisfy the canons of scientific rigor. Quantitative researchers are becoming increasingly aware that some of their data may not be accurate and valid, because some survey respondents may not understand the meaning of questions to which they respond, and because people’s recall of even recent events is often faulty. On the other hand, qualitative researchers have developed better techniques for classifying and analyzing large bodies of descriptive data. It is also increasingly recognized that all data collection - quantitative and qualitative - operates within a cultural context and is affected to some extent by the perceptions and beliefs of investigators and data collectors.

Philosophical distinction. Some researchers and scholars differ about the respective merits of the two approaches largely because of different views about the nature of knowledge and how knowledge is best acquired. Many qualitative researchers argue that there is no objective social reality, and that all knowledge is "constructed" by observers who are the product of traditions, beliefs, and the social and political environment within which they operate. And while quantitative researchers no longer believe that their research methods yield absolute and objective truth, they continue to adhere to the scientific model and seek to develop increasingly sophisticated techniques and statistical tools to improve the measurement of social phenomena. The qualitative approach emphasizes the importance of understanding the context in which events and outcomes occur, whereas quantitative researchers seek to control the context by using random assignment and multivariate analyses. Similarly, qualitative researchers believe that the study of deviant cases provides important insights for the interpretation of findings; quantitative researchers tend to ignore the small number of deviant and extreme cases.

This distinction affects the nature of research designs. According to its most orthodox practitioners, qualitative research does not start with narrowly specified evaluation questions; instead, specific questions are formulated after open-ended field research has been completed (Lofland and Lofland, 1995). This approach may be difficult for program and project evaluators to adopt, since specific questions about the effectiveness of interventions being evaluated are usually expected to guide the evaluation. Some researchers have suggested that a distinction be made between Qualitative and qualitative work: Qualitative work (large Q) refers to methods that eschew prior evaluation questions and hypothesis testing, whereas qualitative work (small q) refers to open-ended data collection methods such as indepth interviews embedded in structured research (Kidder and Fine, 1987). The latter are more likely to meet EHR evaluators' needs.

Practical issues. On the practical level, there are four issues which can affect the choice of method:
• Credibility of findings;
• Staff skills;
• Costs; and
• Time constraints.
Credibility of findings. Evaluations are designed for various audiences, including funding agencies, policymakers in governmental and private agencies, project staff and clients, researchers in academic and applied settings, as well as various other "stakeholders" (individuals and organizations with a stake in the outcome of a project). Experienced evaluators know that they often deal with skeptical audiences or stakeholders who seek to discredit findings that are too critical or uncritical of a project's outcomes. For this reason, the evaluation methodology may be rejected as unsound or weak for a specific case.

The major stakeholders for EHR projects are policymakers within NSF and the federal government, state and local officials, and decisionmakers in the educational community where the project is located. In most cases, decisionmakers at the national level tend to favor quantitative information because these policymakers are accustomed to basing funding decisions on numbers and statistical indicators. On the other hand, many stakeholders in the educational community are often skeptical about statistics and "number crunching" and consider the richer data obtained through qualitative research to be more trustworthy and informative. A particular case in point is the use of traditional test results, a favorite outcome criterion for policymakers, school boards, and parents, but one that teachers and school administrators tend to discount as a poor tool for assessing true student learning.
Staff skills. Qualitative methods, including indepth interviewing, observations, and the use of focus groups, require good staff skills and considerable supervision to yield trustworthy data. Some quantitative research methods can be mastered easily with the help of simple training manuals; this is true of small-scale, self-administered questionnaires, where most questions can be answered by yes/no checkmarks or selecting numbers on a simple scale. Large-scale, complex surveys, however, usually require more skilled personnel to design the instruments and to manage data collection and analysis.

Costs. It is difficult to generalize about the relative costs of the two methods; much depends on the amount of information needed, quality standards followed for the data collection, and the number of cases required for reliability and validity. A short survey based on a small number of cases (25-50) and consisting of a few "easy" questions would be inexpensive, but it also would provide only limited data. Even cheaper would be substituting a focus group session for a subset of the 25-50 respondents; while this method might provide more "interesting" data, those data would be primarily useful for generating new hypotheses to be tested by more appropriate qualitative or quantitative methods. To obtain robust findings, the cost of data collection is bound to be high regardless of method.

Time constraints. Similarly, data complexity and quality affect the time needed for data collection and analysis. Although technological innovations have shortened the time needed to process quantitative data, a good survey requires considerable time to create and pretest questions and to obtain high response rates. However, qualitative methods may be even more time consuming because data collection and data analysis overlap, and the process encourages the exploration of new evaluation questions (see Chapter 4). If insufficient time is allowed for the evaluation, it may be necessary to curtail the amount of data to be collected or to cut short the analytic process, thereby limiting the value of the findings. For evaluations that operate under severe time constraints - for example, where budgetary decisions depend on the findings - the choice of the best method can present a serious dilemma.

In summary, the debate over the merits of qualitative versus quantitative methods is ongoing in the academic community, but when it comes to the choice of methods for conducting project evaluations, a pragmatic strategy has been gaining increased support. Respected practitioners have argued for integrating the two approaches building on their complementary strengths.1 Others have stressed the advantages of linking qualitative and quantitative methods when performing studies and evaluations, showing how the validity and usefulness of findings will benefit (Miles and Huberman, 1994)

http://www.nsf.gov/pubs/1997/nsf97153/chap_1.htm

VAI TRÒ CỦA THỐNG KÊ TRONG NGHIÊN CỨU KHOA HỌC

Quote:

Introduction to Role of Statistics in the Scientific Method
Statistics has a major role to play in all stages of the scientific method. This is because it is involved with the definition and evaluation of hypotheses through the collection and analysis of data. In the paths (a) --> (b) and (e) --> (b) of analytical and inductive reasoning, the methods of descriptive statistics have their role to play. They provide powerful tools for suggesting questions to ask and formulating hypotheses. This is particularly useful in the study of large data sets, especially those routinely collected without specific research purposes; in mind. Such data should also be examined for indications as to the hypothesis or theoretical model underlying the process which produced the data. Data examination may include exploratory techniques such as tabulations, summary descriptions, graphical analysis, and cluster analysis.
Statisticians play an invaluable role in this exploratory stage by working closely with researchers. A basic understanding of the subject area and excellent communication skills are important for the success of this collaboration.
The experimental process, paths (c) --> (d) --> (e), of the scientific method is intimately is involved with many areas of statistics. A description of the statistical methods and thought processes in this part of the scientific method is depicted in Figure 2 and will now be discussed.

After clearly formulating the statistical hypothesis, relevant and valid data are accumulated from historical records, sample surveys or experiments in order to test the given hypotheses and provide indications for possible alternatives. Statistics provides the researcher with an array of methodologies to help in the design of an efficient and cost-effective data collection scheme which also ensures the accuracy, unbiasedness, and quality of the data.

in the area of measurement process the statistician's technical skills are needed. Close collaboration with researchers in the subject area and communication of statistical principles are also crucial.

The principles of quality control may prove to be of valuable assistance during and after data collection. Data ought to be routinely checked for the presence of errors, biases and outliers. The relevance of the data to the hypotheses under study also needs to be continually checked.

Statistics plays a crucial role, in experimentation. Even in the best planned experiments, we cannot control all the factors that affect our observations and we can rarely make. measurements without some noise or error from the measurement process. Hence, we have to make inferences based on imprecise sample data (emphasis by PA). To be of practical use, these uncertain inferences must be accompanied by probability statements expressing the degree of confidence the researcher has in the conclusions. To make certain that such probability statements will be possible, the experiments should be designed in accordance with the principles of statistical experimental design. These principles, together with the statistical hypothesis under study, dictate a statistical model relating the data to the statistical hypothesis through probability theory.

In other words, data have no meaning in themselves; they are meaningful only in relation to a statistical model of the phenomenon being studied. The interpretation of a set of data would be different, depending on what model was thought appropriate. In practice, some basic knowledge of the phenomenon under study is usually available to allow the researcher to specify a plausible statistical model.

http://www.pustaka-deptan.go.id/rkb/knowledgeBank/IPM/stats/stats.htm#Introduction_to_Role_of_Statistics_in_the_Scientific_Method.htm

CÁC ĐẶC ĐIỂM CỦA PHƯƠNG PHÁP XỬ LÝ DỮ LIỆU BẰNG THỐNG KÊ

- Rút ra quy luật thông qua việc quan sát các hiện tượng lập đi lập lại, sử dụng phương pháp quy nạp và kinh nghiệm thực (thực nghiệm)
- Tư duy có hệ thống và khoa học dựa trên các quy luật của số lớn
- Cho phép suy đoán về tổng thể từ một mẫu nhỏ hơn

Thứ Bảy, 27 tháng 12, 2008

Seeing Through Statistics (1): Case Study 3.1

Khi thực hiện một cuộc khảo sát, cần nhớ rằng kết quả khảo sát phụ thuộc rất nhiều vào kỹ thuật và phương pháp khảo sát. Hãy đọc Case sau đây và rút ra kết luận cho chính bạn. Và hãy liên hệ với 7 pitfalls đã nêu trong blog này.

No Opinion of Your Own: Let Politics Decide
(Source: Morin 10-16 April 1995, p. 36)

This is an exellent example of how people will respond to survey questions, even they do not know about the issues, and how the wording of questions can influence responses. In 1995, the Washington Post decided to expand on a 1978 poll taken in Cincinnatti, Ohio, in which people were asked whether they "favoured or opposed repealing the 1975 Public Affairs Act." There was no such act, but about one third of the respondents expressed an opinion about it.

In February 1995, the Washington Post added this fictitious question to its weekly poll of 1000 randomly selected respondents: "Some people say the 1975 Public Affairs Act should be repealed. Do you agree or disagree that it should be repealed?" Almost half of the sample (43%) expressed an opinion, with 24% agreeing that it should be repealed and 19% disagreeing. The Post then tried another trick that produced even more disturbing results. This time, they polled two separate groups of 500 randomly selected adults. The first group was asked: "President Clinton [a Democrat] said that the 1975 Public Affairs Act should be repealed. Do you agree or disagree?" (còn tiếp)

Pitfalls that can be encountered when asking questions in a survey (Seeing Through Statistics - STS)

(STS p38)

Many pitfalls can be encountered when asking questions in a survey or experiment. Here are some of them:

1. Deliberate bias (cố tình thiên vị)
2. Unintentional bias (vô tình thiên lệch)
3. Desire to please (muốn làm hài lòng)
4. Asking the uninformed (hỏi người không biết)
5. Unnecessary complexity (vô cùng rắc rối)
6. Ordering of questions (thứ tự sắp xếp)
7. Confidentiality and anonymity (bí mật khuyết danh)

From: Seeing Through Statistics: Những cảm nhận đầu tiên

Mục này dành để đưa lên các trích dẫn đáng đọc trong sách: Seeing Through Statistics (Jessica M. Utts, 3rd ed., 2005. Thomson, Belmont, CA; Printed in Canada)

Thống kê có lẽ là môn duy nhất đáng học đối với một người có học trong xã hội ngày nay. Vì mọi quyết định đều phải dựa trên thống kê. Thật ra, dù có học thống kê hay không thì mọi người vẫn sử dụng các phương pháp suy luận theo thống kê mỗi khi ra quyết định. Nhưng không học thì chỉ có thể suy luận được đến một mức độ nào đó, còn sau đó thì bắt đầu sai vì thiếu phương pháp. Vì vậy, với thống kê, chúng ta có thể suy luận và ra quyết định đúng đắn hơn TRONG ĐIỀU KIỆN KHÔNG CÓ ĐỦ THÔNG TIN!

Quan trọng như vậy mà ở VN thì chẳng mấy ai được học cho tử tế. Trong đó có tôi. Mặc dù tôi vẫn luôn cố gắng học hỏi, và cố gắng áp dụng một số phương pháp thống kê mà tôi đã từng học được. Và cũng không phải là không biết gì.

Thôi chưa giỏi thì phải tiếp tục học, chứ sao! Cuốn sách Seeing Through Statistics là một cuốn sách mà tôi cho là rất hữu ích, và tôi đang học nó. Học một mình, tự học đấy. Nhưng không đến đâu. Thành ra, phải viết ra, trước hết là cho chính mình, và sau đó là cho ai đó có lang thang đến trang blog này thì cũng có thể được chia sẻ.

Vậy thì, tất cả những người sợ, thích, hoặc ghét statistics, hãy cùng chia sẻ về statistics với tôi trên trang web này nhé!