CCL: Restart CCSD jobs
- From: "Van Dam, Hubertus J"
<HubertusJJ.vanDam{:}pnnl.gov>
- Subject: CCL: Restart CCSD jobs
- Date: Mon, 22 Oct 2012 10:08:35 -0700
Sent to CCL by: "Van Dam, Hubertus J" [HubertusJJ.vanDam|,|pnnl.gov]
Hi Roger,
I think what you are doing is probably alright. The main gotcha that you might
run into is that the number representations might be different on different
machines. However, as in recent years the number of different processor
architectures in common use has decreased that is less of a problem today than
it used to be. Having said that I have trouble pinpointing what bit of output
you are looking at. If you want me to provide some more detailed comments could
you send your input file, please?
Best wishes,
Huub van Dam
Pacific Northwest National Laboratory
Tel: 509-372-6441
-----Original Message-----
> From: owner-chemistry+hubertus.vandam==pnnl.gov]*[ccl.net [mailto:owner-chemistry+hubertus.vandam==pnnl.gov]*[ccl.net]
On Behalf Of Roger Robinson r.robinson*o*imperial.ac.uk
Sent: Monday, October 22, 2012 8:09 AM
To: Van Dam, Hubertus J
Subject: CCL: Restart CCSD jobs
Sent to CCL by: Roger Robinson [r.robinson : imperial.ac.uk] Hi,
I need to restart some CCSD jobs as the machine they are on are being
decommissioned.
The jobs have produced some very large rwf ie between 100GB and 270GB.
I have copied the rwf files back from local sites to the network device.
I notice the RWF havent been rewritten to for several days.
They have all got to the part of the job after the convergence iterations ie all
say something like this.
Iteration Nr. 14
**********************
DD1Dir will call FoFDir 1 times, MxPair= 600
NAB= 300 NAA= 0 NBB= 0 NumPrc= 16.
DE(Corr)= -1.5338361 E(CORR)= -349.45990574 Delta=-1.89D-08
NORM(A)= 0.12459143D+01
Largest amplitude= 3.48D-02
The scratch files have also stopped growing.
Is there still alot of input and output to the scratch files anymore after this
point ?
Would it be possible to carry out the jobs with the scratch file on the network
drive at this point ? I should I move then to new local file systems ?
As a test I restarted one job with the restart file on the network file system.
The job seems to be running ok and more importantly is using 100% of the CPU
allocated which could suggest its not struggling for I/O throughtput ?
Thanks Rogerhttp://www.ccl.net/cgi-bin/ccl/send_ccl_messagehttp-:-//www.ccl.net/chemistry/sub_unsub.shtmlhttp-:-//www.ccl.net/spammers.txt