t test - help interpreting it
Posted in 2004
Topics: General Discussion
Hi all, Actually I am confusig myself in doing this . I have two sets of data as follows: (A) 1 64.125 2 62.061 3 86.084 4 56.109 5 72.141 6 78.076 7 84.082 Mean 96.152 SD 37.729 (B) 1 48.094 2 44.043 3 86.084 4 56.109 5 48.094 6 50.049 7 78.076 Mean 88.189 SD 30.559 (C) 1 16.031 2 18.018 3 0 4 0 5 24.047 6 28.027 7 6.006 Mean 13.16128571 SD 11.32069026 A & B is the actual set of data with their Mean and standard deviation. And C = A - B. I am interested to find in the significance of difference between A & B. So in a way if I can find that wheather C is same as 0 !!. I did t test in that regard (on C). found some result as The t-value for this difference 3.076 p-value= 0.98911; (1-p=0.01088) After this I am confusing interpreting (rejection of null hypothesis). Kindly help!! I hope I have explained the problem clearly enough!! Any help in this regard will be of great value. Thanking you in anticipation .
Young Engineer wrote: > Actually I am confusing myself in doing this. I have two sets of data > as follows: > > (A) > 1 64.125 > 2 62.061 > 3 86.084 > 4 56.109 > 5 72.141 > 6 78.076 > 7 84.082 > > Mean 96.152 This mean is meaningless - since the mean is higher than the highest data value, it is 100% bogus. > SD 37.729 I got: # Count = 7 # Sum(x1) = 5.026780e+02 # Sum(x2) = 3.689223e+04 # Sum(x3) = 2.763107e+06 # Sum(x4) = 2.107957e+08 # Mean = 7.181114e+01 # Std Dev = 1.150612e+01 # Variance = 1.323908e+02 # Skew = -2.356440e-02 # Kurtosis = 1.133173e+00 # Min = 5.610900e+01 # Max = 8.608400e+01 > (B) > > 1 48.094 > 2 44.043 > 3 86.084 > 4 56.109 > 5 48.094 > 6 50.049 > 7 78.076 > > Mean 88.189 > SD 30.559 Likewise, these values are bogus. For this set, I got: # Count = 7 # Sum(x1) = 4.105490e+02 # Sum(x2) = 2.572529e+04 # Sum(x3) = 1.723793e+06 # Sum(x4) = 1.227232e+08 # Mean = 5.864986e+01 # Std Dev = 1.656628e+01 # Variance = 2.744417e+02 # Skew = 6.867854e-01 # Kurtosis = 1.488418e+00 # Min = 4.404300e+01 # Max = 8.608400e+01 (Statistics from pstats.pl - a Perl script by yours truly.) We can discuss sample vs population variances another day. > (C) >[...deleted...] > > A & B is the actual set of data with their Mean and standard > deviation. > And C = A - B. I am interested to find in the significance of > difference between A & B. So in a way if I can find that wheather C is > same as 0 !!. I did t test in that regard (on C). found some result as It is a depressingly long time since I had to do statistics, but I don't think this is the way you do it, is it? For starters, the t-test checks the difference between two samples, and C is only one sample. You have to work with both sets of data. The null hypothesis is that sets A and B are representative samples of a single population. You want to show that this is probably incorrect - so the two samples are different. You calculate the t value from ("Statistical Methods and the Geographer", S Gregory, Longmans, 1963 - no ISBN since they weren't invented until the end of the 60's (and no, I didn't get that book from new; toddlers don't read at undergraduate level, not even me)): t = difference between the means / standard error of the difference t = abs(mean(a) - mean(b)) / sqrt(var(a)/num(a) + var(b)/num(b)) t = abs(71.811 - 58.650) / sqrt(132.391/7 + 274.442/7) = 13.161 / sqrt(18.913 + 39.206) = 13.161 / sqrt(58.119) = 13.161 / 7.624 = 1.726 N = (num(a) - 1) + (num(b) - 1) = 12 A t value of 1.726 wih 12 degrees of freedom is significant at somewhere about 10% probability - the graph I have suggests that the 5% t-value for 12 df is about 2.3, and the 10% value is about 1.8. A quick Google feeling lucky on "student's t test statistics table" leads to: http://www.itl.nist.gov/div898/handbook/eda/section3/eda3672.htm And the table gives the values 1.782 for 0.05 and 2.179 for 0.025, and an explanation of why the 0.05 value corresponds to 10% and 0.025 corresponds to 5%. So, the null hypothesis - that samples A and B come from the same population - is not disproven. You need more data. > After this I am confusing interpreting (rejection of null hypothesis). Your data does not reject the null hypothesis. In the absence of more data, you should accept that they are samples from the same total population. -- Jonathan Leffler #include <disclaimer.h> Email: jleffler@earthlink.net, jleffler@us.ibm.com Guardian of DBD::Informix v2003.04 -- http://dbi.perl.org/