An open API service providing commit metadata for open source projects.

GitHub / sql-machine-learning/elasticdl / commits

Kubernetes-native Deep Learning Framework

SHA Message Author Date Stats
0cbc361a Update README.md (#2549) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 2 years ago
86cd6ff7 Import: This repository is deprecated. (#2547) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
about 3 years ago
d7232265 Bump tensorflow from 2.7.2 to 2.9.3 in /elasticdl_preprocessing (#2543)
Co-authored-by: dependabot[bot] <4****]@u****m>
Signed-off-by: dependabot[bot] <s****t@g****m>, dependabot[bot] <s****t@g****m>
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m>
over 3 years ago
22e7099b Bump tensorflow from 2.5.2 to 2.7.2 in /elasticdl_preprocessing (#2540)
Co-authored-by: dependabot[bot] <4****]@u****m>
Signed-off-by: dependabot[bot] <s****t@g****m>, dependabot[bot] <s****t@g****m>
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m>
almost 4 years ago
f7ded492 Bump tensorflow from 2.1.2 to 2.5.2 in /elasticdl (#2534)
Co-authored-by: dependabot[bot] <4****]@u****m>
Signed-off-by: dependabot[bot] <s****t@g****m>
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m>
over 4 years ago
67955cbc Bump tensorflow from 2.5.1 to 2.5.2 in /elasticdl_preprocessing (#2533)
Co-authored-by: dependabot[bot] <4****]@u****m>
Signed-off-by: dependabot[bot] <s****t@g****m>
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m>
almost 5 years ago
5abcdc37 Fix a spelling error (#2532) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
almost 5 years ago
5a74bd7d A PyTorch tutorial (#2530) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
almost 5 years ago
3558e62a Update ReadMe (#2531) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
almost 5 years ago
a5a6c501 A tutorial for tf.estimator (#2528) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
almost 5 years ago
5217774b Add an example for DeepFM estimator in deepctr (#2526) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
almost 5 years ago
87934a7d Add an example of tf.estimator (#2525) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
almost 5 years ago
ddfa3b7f Improve performance (#2529) DLPerf <8****f@u****m>
Committed by: GitHub <n****y@g****m>
almost 5 years ago
2dcdce65 Bump tensorflow from 2.1.2 to 2.5.1 in /elasticdl_preprocessing (#2523)
Co-authored-by: dependabot[bot] <4****]@u****m>
Signed-off-by: dependabot[bot] <s****t@g****m>
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m>
almost 5 years ago
4edfde51 Design for Dynamic Sharding to Support Reading Original Images (#2341) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
447ac930 Add an example of Tensorflow 1.x with pictures. (#2518) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
ceb636da Fix the docstring to submit a pytorch example (#2517) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
73c5c355 Bump jinja2 from 2.11.2 to 2.11.3 in /elasticdl_client (#2515)
Co-authored-by: dependabot[bot] <4****]@u****m>
Signed-off-by: dependabot[bot] <s****t@g****m>
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
d727d3d8 Set need_elasticdl_job_service=true (#2516) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
96bfc8e2 Remove codes about the training_data (#2512) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
a7e6f950 Master will not add any worker if the current rendezvous hosts become empty a... Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
94029cad Remove the argument custom_training_loop (#2509) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
87664190 Remove the word global because it is redundant (#2511) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
ee6288b0 Develop an API to get hooks for elastic optimizers (#2510) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
4bc8a2ad Set a new version (#2506) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
3eb4051e Develop a reader to read records from recordio files (#2502) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
48d37bd3 Remove spaces after colon (#2507) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
0c11a53d Support to create data shards infinitely. (#2505) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
e9493131 The master does not add the worker host if it is in the next group (#2504) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
1559bf5c Remove the code to report the training data (#2501) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
3d1ddcba Support customize training loop and dataset using TensorFlow 1.x (#2500) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
b2966573 Remove unnecessary Worker param (#2499) junfan.zhang <z****a@g****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
e60551bd Fix the counter if the worker retry to aggregate gradient in ElasticDL (#2497) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
0c00b663 Worker can report training data to the master if using RecordIO (#2494) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
6f8754cd Fix the bug if the host of a new worker is the same as an old worker (#2496) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
f7bee19e Delete the codes to download heart dataset (#2498) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
546dd54a Warp the job command using parentheses. (#2495) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
915acffa upgrade grpcio version to 1.34.1 (#2490) HT <t****c@y****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
40f53921 Develop an API to get training epoch (#2488) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
37c8c8f0 Remove "elasticdl-" prefix to ps/worker pod name (#2489) HT <t****c@y****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
299ef387 Check whether to register hooks according to HOROVOD_ELASTIC (#2487) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
65a2ec0a Create an ElasticImageFolder for PyTorch. (#2486) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
7bdfebf0 Relaunch worker on failure (#2485) HT <t****c@y****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
4c5e7d8a Fix model cuda (#2484) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
df07f8e3 add pod status change log (#2483) HT <t****c@y****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
b035d6b1 Implement the fail fast mechanism of master. (#2480) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
6682003d Add cluster_spec_json in EXCLUDE_PRINT_ARGS (#2479) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
1867af75 Update the version releasing doc. (#2474) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
2d2493ad Add num_minibatches_per_shard param in the factory method. (#2476) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
7c48d8b4 Add num_minibatches_per_shard param in report_training_params (#2473) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
39226f18 Add the argument to shuffle shards. (#2472) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
e29c9d1e A pytorch example to read original images with the custom dataloader (#2415) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
d5687b9c Support shuffling the total dataset. (#2469) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
53689bb6 The event type of PodStateFlow contains ADDED and MODIFED (#2470) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
28ebc85e Add more logs for task_manager. (#2471) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
fa85ded6 Master raise a runtime error if all workers failed (#2468) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
1ffd4e57 Read indices from a shard (#2463) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
3cd10dda Populate the environment variables matched with the input args from master to... brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
be5bf6dc Add the PodStateFlow from PENDING to SUCCEEDED&FAILED. (#2466) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
9ed5b558 Add elasticdl job arguments only when need_elasticdl_job_args=True (#2465) HT <t****c@y****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
e156528a create worker service and set TF_CONFIG env when needed (#2462) HT <t****c@y****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
cad9c23f The worker sends the start and end message to the master. (#2451) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
504ce218 Don't print some arguments. (#2464) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
649dc9ab Enable the worker image configuration using elasticdl_client. (#2461) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
30cd6f3e Add more parameters in the factory method of data_shard_service. (#2460) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
5d4f3c70 Fix the warning message in the master (#2458) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
ca30653f The worker sends the training parameters to the master. (#2448) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
901498b8 Only use cluster_spec or cluster_spec_json (#2457) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
1abda706 Set the default value of need_elasticdl_job_service to False (#2456) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
0b96ef1a The master only exit according to the status of workers (#2447) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
57f8f30e Rearrange worker pod priority (#2444) HT <t****c@y****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
93349cf2 Make data_shard_service to be compatible with various task types. (#2455) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
e8b6c372 Fix the shard name in the unittests of odps (#2454) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
6e6e4511 Add a lock to get task (#2453) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
cdff1e49 Add factory method for data_shard_service and master_client. (#2452) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
074fb8e7 Fix the interval to retry to get the rank (#2450) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
16911254 Set backward passes per step for DistributedOptimizer (#2449) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
61d830ca Raise a runtime error if fail to perform allreduce operation (#2446) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
ab239382 Move grpc_utils.py and log_utils.py to util folder in elasticai_api. (#2443) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
8ed0bb2d Add requirements for elasticai_api package. (#2442) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
2dcc7513 Bump version. (#2441) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
80994514 Move DataShardService and AllReduceControllers from elasticdl package to elas... brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
8d99fdff Set the max seconds to check rendezvous (#2439) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
a0414a4e Separate master RPC service to into master and TrainLoopMaster. (#2438) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
9c901b5b Add an argument to disable the thread to check timeout tasks. (#2434) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
f12e83cb Remove get_model_version method from master_client because Master service has... brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
206dbee0 Move the proto messages about dynamic sharding into elasticai_api.proto. (#2435) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
c3615f44 Move the common constants into elasticai_api and add elasticai_api in the req... brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
b78a60cf Add up the completed global batch count. (#2430) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
39bfa3e9 Keep the name of args in torch optimizer same as TF opt (#2429) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
e18b2085 Conver timout seconds from env to int (#2428) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
193d03bd Use the max completed time of task to check timeout tasks. (#2424) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
cbfdf8ee Master does not exit if there are no worker (#2420) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
8199c979 Initial version of elasticai_api package. (#2422) brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
b3954607 Fix the bug without executing zero_grad (#2426) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
343004fd Check whether to initialize Horovod periodically according to the timeout of ... Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
ba680f87 Use elastic Horovod to run model locally (#2425) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
452ab49b Log the rank and world size when the worker initilize Horvood (#2423) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
7a64aae5 Add an unittest for pytorch DistributedOptimizer (#2421) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago
d2151c69 Check wether the task is a validate training task by task type (#2419) Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m>
over 5 years ago

← Back to repository