From chemistry-request@ccl.net Thu Apr 14 12:26:40 2005
Received: from smtp.goldrush.com (smtp.goldrush.com [206.171.171.11])
	by server.ccl.net (8.13.1/8.13.1) with ESMTP id j3EGQZZF016535
	for <chemistry)at(ccl.net>; Thu, 14 Apr 2005 12:26:35 -0400
Received: from Compaq (x2-04-151.goldrush.com [64.162.10.151])
	by smtp.goldrush.com (8.12.8/8.12.8) with SMTP id j3EGQJC1023826;
	Thu, 14 Apr 2005 09:26:27 -0700
From: "Steve Bowlus" <chezbowlus)at(goldrush.com>
To: <lhhu)at(yangtze.hku.hk>, <chemistry)at(ccl.net>
Subject: RE: Correlation Coefficient
Date: Thu, 14 Apr 2005 09:26:20 -0700
Message-ID: <OLEDKJNBCEKJDJFPILMJIEHKCAAA.chezbowlus)at(goldrush.com>
MIME-Version: 1.0
Content-Type: text/plain;
	charset="Windows-1252"
Content-Transfer-Encoding: 7bit
X-Priority: 3 (Normal)
X-MSMail-Priority: Normal
X-Mailer: Microsoft Outlook IMO, Build 9.0.2416 (9.0.2910.0)
Importance: Normal
In-Reply-To: <Pine.LNX.4.61.0504141219570.2286)at(yangtze.hku.hk>
X-MimeOLE: Produced By Microsoft MimeOLE V6.00.2900.2180
X-MailScanner: Found to be clean
X-MailScanner-SpamCheck: 
X-MailScanner-From: chezbowlus)at(goldrush.com
X-Spam-Status: No, score=0.1 required=5.0 tests=DNS_FROM_AHBL_RHSBL 
	autolearn=failed version=3.0.1
X-Spam-Checker-Version: SpamAssassin 3.0.1 (2004-10-22) on server.ccl.net

Cross-validated r-sq (or q-sq) provides a measure of the homogeneity of the
data set.  Depending on how the model is developed, it reveals whether there
are influential points in the data set which are unduly weighted in the
model.  In this regard, q-sq replaces the influence statistics common in
univariate statistics.

By "reliable" are you actually meaning to ask whether predictions made by
the model will be accurate (i.e. useful)?  One must remember that 1)
predictions are associated with a confidence interval, and 2) the new
compound must lie in the prediction space of the model (not be an outlier
wrt any of the descriptors).  So prediction made with a large CI may be
statistically correct, but practically useless ("unreliable"?).  Bottom
line, the model should have high r-sq, high cross-validated r-sq, and a
reasonably tight standard error of prediction.

What is useful is of course dependent on the situation.  Small, noisy data
sets at the beginning of a project, where the intent of the model may be to
determine areas or descriptors for exploration, might use r-sq on the order
of 0.7 and q-sq as low as 0.3.  As the dataset grows and becomes more
homogeneous in the descriptor space, r-sq and q-sq should both increase and
converge on the same value, while the SE shrinks to a (practical) limit of
the accuracy of the (bio)assay.

One or more journal editors may weigh in with their favorite cutoffs, but I
am not aware there is any magic number to assure the "reliability" of a
model.

Steve Bowlus




-----Original Message-----
From: Computational Chemistry List [mailto:chemistry-request)at(ccl.net]On
Behalf Of lhhu)at(yangtze.hku.hk
Sent: Wednesday, April 13, 2005 9:49 PM
To: chemistry)at(ccl.net
Subject: CCL:Correlation Coefficient



Dear All,

I have a question about Correlation Coefficient of cross-validation. Can
anyone tell me how much the value of Correlation Coefficient indicate
the result is reliable? Certainly I know the larger the value, the more
reliable the result. I want to know the lowest limit of that value. I read
some papers that used the Correlation Coefficient value, some values are
lower than 0.8, is it OK for present them in journals? Or it depends on
different situation or topic?

Thanks a lot to you all,

Holly





-= This is automatically added to each message by the mailing script =-
To send e-mail to subscribers of CCL put the string CCL: on your Subject:
line
and send your message to:  CHEMISTRY)at(ccl.net

Send your subscription/unsubscription requests to: CHEMISTRY-REQUEST)at(ccl.net
HOME Page: http://www.ccl.net   | Jobs Page: http://www.ccl.net/jobs

If your is mail bouncing from ccl.net domain due to spam filters, please
use the Web based form from CCL Home Page
-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+










