From owner-chemistry@ccl.net Mon Oct 22 11:47:01 2012 From: "Roger Robinson r.robinson*o*imperial.ac.uk" To: CCL Subject: CCL: Restart CCSD jobs Message-Id: <-47784-121022111052-17690-5c75bLucz3YvDPJO1fT9qA_+_server.ccl.net> X-Original-From: Roger Robinson Content-Transfer-Encoding: 7bit Content-Type: text/plain; charset=ISO-8859-1; format=flowed Date: Mon, 22 Oct 2012 16:08:44 +0100 MIME-Version: 1.0 Sent to CCL by: Roger Robinson [r.robinson : imperial.ac.uk] Hi, I need to restart some CCSD jobs as the machine they are on are being decommissioned. The jobs have produced some very large rwf ie between 100GB and 270GB. I have copied the rwf files back from local sites to the network device. I notice the RWF havent been rewritten to for several days. They have all got to the part of the job after the convergence iterations ie all say something like this. Iteration Nr. 14 ********************** DD1Dir will call FoFDir 1 times, MxPair= 600 NAB= 300 NAA= 0 NBB= 0 NumPrc= 16. DE(Corr)= -1.5338361 E(CORR)= -349.45990574 Delta=-1.89D-08 NORM(A)= 0.12459143D+01 Largest amplitude= 3.48D-02 The scratch files have also stopped growing. Is there still alot of input and output to the scratch files anymore after this point ? Would it be possible to carry out the jobs with the scratch file on the network drive at this point ? I should I move then to new local file systems ? As a test I restarted one job with the restart file on the network file system. The job seems to be running ok and more importantly is using 100% of the CPU allocated which could suggest its not struggling for I/O throughtput ? Thanks Roger