# Job composer and star-ccm+

**URL:** <https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238>\
**Category:** Get Help\
**Tags:** ondemand2, question\
**Created:** [August 17, 2022, 7:22pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238 "2022-08-17T19:22:20Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![jesse.waters](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@jesse.waters](https://discourse.openondemand.org/u/jesse.waters)\
**Post date:** [August 17, 2022, 7:22pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238/1 "2022-08-17T19:22:20Z")

</div>

Have a user submitting a starccm job via the “job composer”

From terminal/commnad line, sbatch job script works as expected .  
Take same script and try to submit it from the job composer and job errors out.

An ORTE daemon has unexpectedly failed after launch and before  
communicating back to mpirun. This could be caused by a number  
of factors, including an inability to create a connection back  
to mpirun due to a lack of common network interfaces and/or no  
route found between them. Please check network connectivity  
(including firewalls and network routing requirements).

This error relates to openmpi that starccm uses by default. Changing starccm to -mpi intel, and submit via job composer it works.

Environment shell appears to be bash all the way through. Any suggestion on where this is getting broken?

---

<div class="post-metadata">

**Author:** ![travert](https://sea1.discourse-cdn.com/flex015/user_avatar/discourse.openondemand.org/travert/32/3074_2.png) [@travert](https://discourse.openondemand.org/u/travert)\
**Post date:** [August 18, 2022, 4:23pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238/2 "2022-08-18T16:23:52Z")

</div>

Sorry for the issue. Reading what you have here I wonder if the fix is similar from a previous user’s issue:

> [@Slurm job fails in OOD but works at CLI](https://discourse.openondemand.org/t/slurm-job-fails-in-ood-but-works-at-cli/935/2):
>
> strace maybe? When you say you can run it from the cli, I’d ask if it’s the same host? The cli from the same server as OOD or some other login host? This could be your discrepancy, there could actually be some networking issue there. It could also be the environment. OOD defaults to SBATCH\_EXPORT=NONE so you can load a brand new environment. I would say add an env statement to see if there’s something you’re missing (some LD\_LIBRARY\_PATH missing?).

The solution seemed to be:

> [@Slurm job fails in OOD but works at CLI](https://discourse.openondemand.org/t/slurm-job-fails-in-ood-but-works-at-cli/935/3):
>
> I’ll bet it’s due to OOD defaulting to `SBATCH_EXPORT=NONE`, I believe that has issues with parallelism. I wonder if adding `#SBATCH --export=ALL` directive to the script overrides this?

Which might be why the environment is not what is expected? What happens when you add that to the line and submit in the job composer?

---

<div class="post-metadata">

**Author:** ![jesse.waters](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@jesse.waters](https://discourse.openondemand.org/u/jesse.waters)\
**Post date:** [August 18, 2022, 5:48pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238/3 "2022-08-18T17:48:48Z")

</div>

Thanks for the response. I tried adding #SBATCH --export=ALL, gives the same error.  
It has to be environment, maybe source in /etc/profile and /etc/bash and see

---

<div class="post-metadata">

**Author:** ![jesse.waters](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@jesse.waters](https://discourse.openondemand.org/u/jesse.waters)\
**Post date:** [August 18, 2022, 6:03pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238/4 "2022-08-18T18:03:20Z")

</div>

After adding `#SBATCH --export=ALL`, job is still submited with SLURM\_EXPORT\_ENV=NONE

What code does the submital?  
/var/www/ood/apps/sys/myjobs

---

<div class="post-metadata">

**Author:** ![travert](https://sea1.discourse-cdn.com/flex015/user_avatar/discourse.openondemand.org/travert/32/3074_2.png) [@travert](https://discourse.openondemand.org/u/travert)\
**Post date:** [August 18, 2022, 6:31pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238/5 "2022-08-18T18:31:35Z")

</div>

It would be something in the `ood_core` that I need to check to understand this, and the exact place where that is being set is here:

> <https://github.com/OSC/ood_core/blob/d577966f44dee12dbd3e419ec806efedb975014f/lib/ood_core/job/adapters/slurm.rb#L718>

So it is odd that the `--export=ALL` won’t work given what I see there. What version of `ood` are you on? Would you be able to post the script or the relevant portions?

---

<div class="post-metadata">

**Author:** ![jesse.waters](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@jesse.waters](https://discourse.openondemand.org/u/jesse.waters)\
**Post date:** [August 18, 2022, 7:22pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238/6 "2022-08-18T19:22:15Z")

</div>

was running on 2.0.20, just upgraded 2.0.28 to see if it was addressed (worth a try, but still same issue).

```auto
#!/bin/bash

#SBATCH --export=ALL
#SBATCH --get-user-env

# Clerical tracking information
#SBATCH --job-name="Test"
#SBATCH --comment="STAR-CCM+ Solution for JOBNAME"
#SBATCH --account="no-code"

#SBATCH --nodes=2 # number of nodes
#SBATCH --ntasks=8 # total number of processor cores
#SBATCH --partition=ondemand # Cluster rack number

#source /etc/profile
#source ~/.bash_profile
#source ~/.bashrc

# Load modules
#module purge
#module avail
#module load modules
#module use /opt/Software/corvid/.modulefiles
#module load star-ccm/16.06.008-R8
#module avail

### VARS available from cli not included on JC
export SSH_CLIENT=
export CHROOTDOR=/var/chroots/sl7
export I_MPI_PMI_LIBRARY=/usr/lib64/libpmi.so
export LANGUAGE_TERRITORY=en_US
export SSH_TTY=/dev/pts/0
export VNFSROOT=sl7
export XMODIFIERS=@im=none
export LANG=en_US.UTF-8
export TERM=linux
export SELINUX_ROLE_REQUESTED=

# Log output
printf "Job debugging info:\n"
echo "----"
printf "Starting starccm+ simulation run at "
date
echo "----"

export SIM_TITLE="DuctTestRun12Copy.sim"

# Create machine file and submit
MACHINEFILE=$(generate_pbs_nodefile)

env
cd /home/corvid/jwaters/Documents/slurm/Star_RUN

### Prefer openmip
#/opt/Software/corvid/star-ccm/16.06.008-R8/STAR-CCM+16.06.008-R*/star/bin/starccm+ -power -machinefile "$MACHINEFILE" -np $SLURM_NTASKS -batch "$SIM_TITLE"
/opt/Software/corvid/star-ccm/16.06.008-R8/STAR-CCM+16.06.008-R*/star/bin/starccm+ -power -mpi "openmpi" -machinefile "$MACHINEFILE" -np $SLURM_NTASKS -batch "$SIM_TITLE"

### Works
/opt/Software/corvid/star-ccm/16.06.008-R8/STAR-CCM+16.06.008-R*/star/bin/starccm+ -power -mpi intel -machinefile "$MACHINEFILE" -np $SLURM_NTASKS -batch "$SIM_TITLE"

```

---

<div class="post-metadata">

**Author:** ![jesse.waters](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@jesse.waters](https://discourse.openondemand.org/u/jesse.waters)\
**Post date:** [August 18, 2022, 7:42pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238/7 "2022-08-18T19:42:54Z")

</div>

Possible similar issue from 2 yrs ago

> [@Interactive app environment in 1.7](https://discourse.openondemand.org/t/interactive-app-environment-in-1-7/878/19):
>
> This has been promoted to stable, so you don’t need to pull it off our latest repo anymore, you can get it from the regular 1.7 repo. Sites that have implemented this fix should not have issues updating directly. There should be no issue sourcing a file a second time (if it gets sourced at all, with the if block there).

---

<div class="post-metadata">

**Author:** ![travert](https://sea1.discourse-cdn.com/flex015/user_avatar/discourse.openondemand.org/travert/32/3074_2.png) [@travert](https://discourse.openondemand.org/u/travert)\
**Post date:** [August 18, 2022, 8:31pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238/8 "2022-08-18T20:31:06Z")

</div>

Yeah looking into it looks like this issue has a fix currently in place for what you need:

> <https://github.com/OSC/ondemand/pull/1847/files>
>
> allow for job composer to copy environment so things like \`srun\` in Slurm work o…f out the box.
> 
> To test, I just copied the default job and began to edit it
> \* allocate more cores through \`#SBATCH --ntasks=5\` comment
> \* add a \`srun hostname\` to the file so it uses srun
> \* run with copy environment checked (this will work)
> \* run without copy environment checked (this will complain about srun)
> 
> 
> 
> ┆Issue is synchronized with this \[Asana task\](https://app.asana.com/0/1201735133575781/1201863073369921) by \[Unito\](https://www.unito.io)

Rebuilding the app off `master` gives the option in the “Job Options” to copy the environment, but it is not in the `2.0` release branch and as of now there are no current plans to back port this, though I am going to add this to the `2.1` milestone.

---

<div class="post-metadata">

**Author:** ![travert](https://sea1.discourse-cdn.com/flex015/user_avatar/discourse.openondemand.org/travert/32/3074_2.png) [@travert](https://discourse.openondemand.org/u/travert)\
**Post date:** [August 19, 2022, 2:55pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238/9 "2022-08-19T14:55:51Z")

</div>

I wanted to update that this will be in `2.1` but until then, I had another idea. Are the jobs for the jobs composer app all going to their own cluster, and using there own `clusters.d` file? If so, what about trying to set the `env` using `bin_overrides` in the `cluster.d/job_composer_cluster.yml` config file:

[https://osc.github.io/ood-documentation/latest/installation/cluster-config-schema.html#bin-overrides](https://osc.github.io/ood-documentation/latest/installation/cluster-config-schema.html#bin-overrides)

And the needed Adapter:

[https://osc.github.io/ood-documentation/latest/installation/resource-manager/slurm.html](https://osc.github.io/ood-documentation/latest/installation/resource-manager/slurm.html)

Let me know if this looks like it might work or if you have any questions around it.

---

<div class="post-metadata">

**Author:** ![jesse.waters](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@jesse.waters](https://discourse.openondemand.org/u/jesse.waters)\
**Post date:** [August 19, 2022, 3:52pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238/10 "2022-08-19T15:52:04Z")

</div>

Able to get the starccm to run.

Work around is to prefix your command w/SLURM\_EXPORT\_ENV=ALL  
Also we have similar issur with srun, use srun --export=ALL

SLURM\_EXPORT\_ENV=ALL /opt/Software/corvid/star-ccm/16.06.008-R8/STAR-CCM+16.06.008-R\*/star/bin/starccm+ -power -mpi openmpi -machinefile “$MACHINEFILE” -np $SLURM\_NTASKS -batch “$SIM\_TITLE”

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex015/uploads/osc/original/2X/b/bae70bd0ed39a3ae769c2108155f4cb3e9da8385.png) [@system](https://discourse.openondemand.org/u/system)\
**Post date:** [February 15, 2023, 3:52pm UTC](https://discourse.openondemand.org/t/job-composer-and-star-ccm/2238/11 "2023-02-15T15:52:53Z")

</div>

This topic was automatically closed 180 days after the last reply. New replies are no longer allowed.
