Overview
The master publishes a REST style API on port 5678.
To submit a job you just post to the /jobs endpoint.
Assuming your job is in a file called job.json' it's as simple as
curl -XPOST "$YOUR_MASTER_IP:5678/v1/jobs" -d@job.json
This will return the job_id (for access to the original job posted) and the job execution context id (the running instance of a job) which can then be used to manage the job. This will also start the job.
{
"job_id": "5a50580c-4a50-48d9-80f8-ac70a00f3dbd",
}
Job Control
Please check the api docs at the bottom for a comprehensive in-depth list of all api's. What is listed here is just a small brief of only a few api's
Job status
This will retrieve the job configuration including '_status' which indicates the execution status of the job.
curl "$YOUR_MASTER_IP:5678/v1/jobs/${job_id}/ex"
Stopping a job
Stopping a job stops all execution and frees the workers being consumed by the job on the cluster.
curl -XPOST "$YOUR_MASTER_IP:5678/v1/jobs/${job_id}/_stop"
Starting a job
Posting a new job will automatically start the job. If the job already exists then using the endpoint below will start a new one.
curl -XPOST "$YOUR_MASTER_IP:5678/v1/jobs/${job_id}/_start"
Starting a job with recover will attempt to replay any failed slices from previous runs and will then pickup where it left off. If there are no failed slices the job will simply resume from where it was stopped.
curl -XPOST "$YOUR_MASTER_IP:5678/v1/jobs/${job_id}/_recover"
Pausing a job
Pausing a job will stop execution of the job on the cluster but will not release the workers being used by the job. It simply pauses the execution and stops allocating work to the workers. Workers will complete the work they're doing then just sit idle until the job is resumed.
curl -XPOST "$YOUR_MASTER_IP:5678/v1/jobs/${job_id}/_pause"
Resuming a job
Resuming a job restarts the execution and the allocation of slices to workers.
curl -XPOST "$YOUR_MASTER_IP:5678/v1/jobs/${job_id}/_resume"
Setting a job active or inactive
A Teraslice job can be marked active or inactive. By default, a job is
active (if the job object does not have the active property, it is assumed
to be active).
curl -XPOST "$YOUR_MASTER_IP:5678/v1/jobs/${job_id}/_active"
curl -XPOST "$YOUR_MASTER_IP:5678/v1/jobs/${job_id}/_inactive"
Changing job settings at runtime
You can dynamically change certain job settings without restarting the job. Currently log_level is supported.
curl -XPOST "$YOUR_MASTER_IP:5678/v1/jobs/${job_id}/_settings?log_level=debug"
Valid log levels are: trace, debug, info, warn, error, fatal.
Viewing Slicer statistics for a job
This provides information related to the execution controller and can be useful in monitoring and optimizing the execution of the job.
curl "$YOUR_MASTER_IP:5678/v1/jobs/${job_id}/controller"
Tapping a slice
This captures the records each operation produces for the next slice a worker
processes, which can be useful when debugging a job's operations. size limits
the records returned per operation (default 10, all for every record).
Warning
Records are captured when each operation completes, so a job with large slices and several operations can produce a very large response. This response will consume memory on the worker, execution controller and the master, possibly causing any or all of them to run out of memory. Use the size query option with extreme caution.
curl "$YOUR_MASTER_IP:5678/v1/jobs/${job_id}/tap?size=5"
If you have the execution id rather than the job id, you can tap the running execution directly:
curl "$YOUR_MASTER_IP:5678/v1/ex/${ex_id}/tap?size=5"
The request waits for a new slice to finish, so it can take a while on slow jobs. See the JSON API docs for the response format and errors.
Viewing cluster state
This will show you all the connected workers and the tasks that are currently assigned to them.
curl "$YOUR_MASTER_IP:5678/v1/cluster/state"