Mailing List Archives Public Access	UW Madison Computer Sciences Department Computer Systems Lab

[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Re: [Condor-users] STARTD-based memory limit

Date: Mon, 06 Jun 2011 13:40:16 -0400
From: Matthew Farrellee <matt@xxxxxxxxxx>
Subject: Re: [Condor-users] STARTD-based memory limit

On 06/02/2011 10:15 AM, Steven Timm wrote:


In my cluster I have been using a schedd-based method of
killing jobs that are using too much memory.

[root@fcdf1x1 local]# condor_config_val SYSTEM_PERIODIC_REMOVE
(NumJobStarts > 10) || (ImageSize>=2500000) || (JobRunCount>=1 &&
JobStatus==1 && ImageSize>=1000000)

But this has two weaknesses

One is that sometimes it can take
the shadow a long time to send the high memory value back to
the schedd so the schedd can act, and in the meantime the job grows
too fast and sucks up all ram on the node and starts killing other
processes.

The second one is that I have a diverse pool of nodes and
would like jobs running on the nodes with bigger memory to use it if
it is there.

So is there a way to evict jobs that use, (ImageSize*2>Memory)?
would you use the KILL or the PREEMPT function?

Steve Timm

Often policy evaluation is delegated to the Shadow. Maybe it's a bugthat SYSTEM_PERIODIC_REMOVE is not.


Best,


matt

Follow-Ups:
- Re: [Condor-users] STARTD-based memory limit
  - From: Dan Bradley

References:
- [Condor-users] STARTD-based memory limit
  - From: Steven Timm

Prev by Date: [Condor-users] DIDC 2011 (June 8th, San Jose) - Call for Participation
Next by Date: Re: [Condor-users] User based authentication in Condor
Previous by thread: Re: [Condor-users] STARTD-based memory limit
Next by thread: Re: [Condor-users] STARTD-based memory limit
Index(es):
- Date
- Thread

Mailing List Archives

Public Access

Re: [Condor-users] STARTD-based memory limit