GitHub / sql-machine-learning/elasticdl / commits
Kubernetes-native Deep Learning Framework
| SHA | Message | Author | Date | Stats |
|---|---|---|---|---|
| 0cbc361a | Update README.md (#2549) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 2 years ago | |
| 86cd6ff7 | Import: This repository is deprecated. (#2547) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
about 3 years ago | |
| d7232265 |
Bump tensorflow from 2.7.2 to 2.9.3 in /elasticdl_preprocessing (#2543)
Co-authored-by: dependabot[bot] <4****]@u****m> Signed-off-by: dependabot[bot] <s****t@g****m>, dependabot[bot] <s****t@g****m> |
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m> |
over 3 years ago | |
| 22e7099b |
Bump tensorflow from 2.5.2 to 2.7.2 in /elasticdl_preprocessing (#2540)
Co-authored-by: dependabot[bot] <4****]@u****m> Signed-off-by: dependabot[bot] <s****t@g****m>, dependabot[bot] <s****t@g****m> |
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m> |
almost 4 years ago | |
| f7ded492 |
Bump tensorflow from 2.1.2 to 2.5.2 in /elasticdl (#2534)
Co-authored-by: dependabot[bot] <4****]@u****m> Signed-off-by: dependabot[bot] <s****t@g****m> |
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m> |
over 4 years ago | |
| 67955cbc |
Bump tensorflow from 2.5.1 to 2.5.2 in /elasticdl_preprocessing (#2533)
Co-authored-by: dependabot[bot] <4****]@u****m> Signed-off-by: dependabot[bot] <s****t@g****m> |
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m> |
almost 5 years ago | |
| 5abcdc37 | Fix a spelling error (#2532) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
almost 5 years ago | |
| 5a74bd7d | A PyTorch tutorial (#2530) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
almost 5 years ago | |
| 3558e62a | Update ReadMe (#2531) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
almost 5 years ago | |
| a5a6c501 | A tutorial for tf.estimator (#2528) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
almost 5 years ago | |
| 5217774b | Add an example for DeepFM estimator in deepctr (#2526) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
almost 5 years ago | |
| 87934a7d | Add an example of tf.estimator (#2525) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
almost 5 years ago | |
| ddfa3b7f | Improve performance (#2529) |
DLPerf <8****f@u****m>
Committed by: GitHub <n****y@g****m> |
almost 5 years ago | |
| 2dcdce65 |
Bump tensorflow from 2.1.2 to 2.5.1 in /elasticdl_preprocessing (#2523)
Co-authored-by: dependabot[bot] <4****]@u****m> Signed-off-by: dependabot[bot] <s****t@g****m> |
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m> |
almost 5 years ago | |
| 4edfde51 | Design for Dynamic Sharding to Support Reading Original Images (#2341) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 447ac930 | Add an example of Tensorflow 1.x with pictures. (#2518) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| ceb636da | Fix the docstring to submit a pytorch example (#2517) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 73c5c355 |
Bump jinja2 from 2.11.2 to 2.11.3 in /elasticdl_client (#2515)
Co-authored-by: dependabot[bot] <4****]@u****m> Signed-off-by: dependabot[bot] <s****t@g****m> |
dependabot[bot] <4****]@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| d727d3d8 | Set need_elasticdl_job_service=true (#2516) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 96bfc8e2 | Remove codes about the training_data (#2512) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| a7e6f950 | Master will not add any worker if the current rendezvous hosts become empty a... |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 94029cad | Remove the argument custom_training_loop (#2509) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 87664190 | Remove the word global because it is redundant (#2511) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| ee6288b0 | Develop an API to get hooks for elastic optimizers (#2510) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 4bc8a2ad | Set a new version (#2506) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 3eb4051e | Develop a reader to read records from recordio files (#2502) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 48d37bd3 | Remove spaces after colon (#2507) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 0c11a53d | Support to create data shards infinitely. (#2505) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| e9493131 | The master does not add the worker host if it is in the next group (#2504) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 1559bf5c | Remove the code to report the training data (#2501) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 3d1ddcba | Support customize training loop and dataset using TensorFlow 1.x (#2500) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| b2966573 | Remove unnecessary Worker param (#2499) |
junfan.zhang <z****a@g****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| e60551bd | Fix the counter if the worker retry to aggregate gradient in ElasticDL (#2497) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 0c00b663 | Worker can report training data to the master if using RecordIO (#2494) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 6f8754cd | Fix the bug if the host of a new worker is the same as an old worker (#2496) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| f7bee19e | Delete the codes to download heart dataset (#2498) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 546dd54a | Warp the job command using parentheses. (#2495) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 915acffa | upgrade grpcio version to 1.34.1 (#2490) |
HT <t****c@y****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 40f53921 | Develop an API to get training epoch (#2488) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 37c8c8f0 | Remove "elasticdl-" prefix to ps/worker pod name (#2489) |
HT <t****c@y****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 299ef387 | Check whether to register hooks according to HOROVOD_ELASTIC (#2487) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 65a2ec0a | Create an ElasticImageFolder for PyTorch. (#2486) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 7bdfebf0 | Relaunch worker on failure (#2485) |
HT <t****c@y****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 4c5e7d8a | Fix model cuda (#2484) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| df07f8e3 | add pod status change log (#2483) |
HT <t****c@y****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| b035d6b1 | Implement the fail fast mechanism of master. (#2480) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 6682003d | Add cluster_spec_json in EXCLUDE_PRINT_ARGS (#2479) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 1867af75 | Update the version releasing doc. (#2474) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 2d2493ad | Add num_minibatches_per_shard param in the factory method. (#2476) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 7c48d8b4 | Add num_minibatches_per_shard param in report_training_params (#2473) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 39226f18 | Add the argument to shuffle shards. (#2472) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| e29c9d1e | A pytorch example to read original images with the custom dataloader (#2415) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| d5687b9c | Support shuffling the total dataset. (#2469) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 53689bb6 | The event type of PodStateFlow contains ADDED and MODIFED (#2470) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 28ebc85e | Add more logs for task_manager. (#2471) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| fa85ded6 | Master raise a runtime error if all workers failed (#2468) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 1ffd4e57 | Read indices from a shard (#2463) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 3cd10dda | Populate the environment variables matched with the input args from master to... |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| be5bf6dc | Add the PodStateFlow from PENDING to SUCCEEDED&FAILED. (#2466) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 9ed5b558 | Add elasticdl job arguments only when need_elasticdl_job_args=True (#2465) |
HT <t****c@y****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| e156528a | create worker service and set TF_CONFIG env when needed (#2462) |
HT <t****c@y****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| cad9c23f | The worker sends the start and end message to the master. (#2451) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 504ce218 | Don't print some arguments. (#2464) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 649dc9ab | Enable the worker image configuration using elasticdl_client. (#2461) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 30cd6f3e | Add more parameters in the factory method of data_shard_service. (#2460) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 5d4f3c70 | Fix the warning message in the master (#2458) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| ca30653f | The worker sends the training parameters to the master. (#2448) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 901498b8 | Only use cluster_spec or cluster_spec_json (#2457) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 1abda706 | Set the default value of need_elasticdl_job_service to False (#2456) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 0b96ef1a | The master only exit according to the status of workers (#2447) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 57f8f30e | Rearrange worker pod priority (#2444) |
HT <t****c@y****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 93349cf2 | Make data_shard_service to be compatible with various task types. (#2455) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| e8b6c372 | Fix the shard name in the unittests of odps (#2454) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 6e6e4511 | Add a lock to get task (#2453) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| cdff1e49 | Add factory method for data_shard_service and master_client. (#2452) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 074fb8e7 | Fix the interval to retry to get the rank (#2450) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 16911254 | Set backward passes per step for DistributedOptimizer (#2449) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 61d830ca | Raise a runtime error if fail to perform allreduce operation (#2446) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| ab239382 | Move grpc_utils.py and log_utils.py to util folder in elasticai_api. (#2443) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 8ed0bb2d | Add requirements for elasticai_api package. (#2442) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 2dcc7513 | Bump version. (#2441) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 80994514 | Move DataShardService and AllReduceControllers from elasticdl package to elas... |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 8d99fdff | Set the max seconds to check rendezvous (#2439) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| a0414a4e | Separate master RPC service to into master and TrainLoopMaster. (#2438) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 9c901b5b | Add an argument to disable the thread to check timeout tasks. (#2434) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| f12e83cb | Remove get_model_version method from master_client because Master service has... |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 206dbee0 | Move the proto messages about dynamic sharding into elasticai_api.proto. (#2435) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| c3615f44 | Move the common constants into elasticai_api and add elasticai_api in the req... |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| b78a60cf | Add up the completed global batch count. (#2430) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 39bfa3e9 | Keep the name of args in torch optimizer same as TF opt (#2429) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| e18b2085 | Conver timout seconds from env to int (#2428) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 193d03bd | Use the max completed time of task to check timeout tasks. (#2424) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| cbfdf8ee | Master does not exit if there are no worker (#2420) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 8199c979 | Initial version of elasticai_api package. (#2422) |
brightcoder01 <5****1@u****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| b3954607 | Fix the bug without executing zero_grad (#2426) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 343004fd | Check whether to initialize Horovod periodically according to the timeout of ... |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| ba680f87 | Use elastic Horovod to run model locally (#2425) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 452ab49b | Log the rank and world size when the worker initilize Horvood (#2423) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| 7a64aae5 | Add an unittest for pytorch DistributedOptimizer (#2421) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago | |
| d2151c69 | Check wether the task is a validate training task by task type (#2419) |
Qinlong Wang <W****1@o****m>
Committed by: GitHub <n****y@g****m> |
over 5 years ago |