jobmanager异常

classic Classic list List threaded Threaded
2 messages Options
Reply | Threaded
Open this post in threaded view
|

jobmanager异常

18500348251@163.com
请教大家一个问题:

    flink1.8.0 on yarn 程序运行一段时间报如下错误,导致 The heartbeat of TaskManager with id container_1572430463280_50994_01_000004 timed out. 最终程序重启。

    各位有没有碰到类似的问题,有什么解决方式吗?

jobmanager.log

2020-08-17 02:53:21,593 ERROR akka.remote.Remoting                                          - Association to [akka.tcp://flink@${HOSTNAME}:36968] with UID [19
99537927] irrecoverably failed. Quarantining address.
java.util.concurrent.TimeoutException: Remote system has been silent for too long. (more than 48.0 hours)
        at akka.remote.ReliableDeliverySupervisor$$anonfun$idle$1.applyOrElse(Endpoint.scala:375)
        at akka.actor.Actor$class.aroundReceive(Actor.scala:502)
        at akka.remote.ReliableDeliverySupervisor.aroundReceive(Endpoint.scala:203)
        at akka.actor.ActorCell.receiveMessage(ActorCell.scala:526)
        at akka.actor.ActorCell.invoke(ActorCell.scala:495)
        at akka.dispatch.Mailbox.processMailbox(Mailbox.scala:257)
        at akka.dispatch.Mailbox.run(Mailbox.scala:224)
        at akka.dispatch.Mailbox.exec(Mailbox.scala:234)
        at scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)
        at scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)
        at scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)
        at scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)





[hidden email]
Reply | Threaded
Open this post in threaded view
|

Re: jobmanager异常

Congxian Qiu
Hi
   container timeout 可以看下是不是 GC 的原因,看一下超时的这个 container
container_1572430463280_50994_01_000004 的之前的 GC 情况
Best,
Congxian


[hidden email] <[hidden email]> 于2020年8月17日周一 上午11:57写道:

> 请教大家一个问题:
>
>     flink1.8.0 on yarn 程序运行一段时间报如下错误,导致 The heartbeat of TaskManager with
> id container_1572430463280_50994_01_000004 timed out. 最终程序重启。
>
>     各位有没有碰到类似的问题,有什么解决方式吗?
>
> jobmanager.log
>
> 2020-08-17 02:53:21,593 ERROR akka.remote.Remoting
>                   - Association to [akka.tcp://flink@${HOSTNAME}:36968]
> with UID [19
> 99537927] irrecoverably failed. Quarantining address.
> java.util.concurrent.TimeoutException: Remote system has been silent for
> too long. (more than 48.0 hours)
>         at
> akka.remote.ReliableDeliverySupervisor$$anonfun$idle$1.applyOrElse(Endpoint.scala:375)
>         at akka.actor.Actor$class.aroundReceive(Actor.scala:502)
>         at
> akka.remote.ReliableDeliverySupervisor.aroundReceive(Endpoint.scala:203)
>         at akka.actor.ActorCell.receiveMessage(ActorCell.scala:526)
>         at akka.actor.ActorCell.invoke(ActorCell.scala:495)
>         at akka.dispatch.Mailbox.processMailbox(Mailbox.scala:257)
>         at akka.dispatch.Mailbox.run(Mailbox.scala:224)
>         at akka.dispatch.Mailbox.exec(Mailbox.scala:234)
>         at
> scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)
>         at
> scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)
>         at
> scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)
>         at
> scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)
>
>
>
>
>
> [hidden email]
>